Describe the bug
With a custom Chat Completions model (github.copilot.chat.customOAIModels, "thinking": true), VS Code parses streamed reasoning_content but never sends it back on later requests.
OpenAIEndpoint.getCompletionsCallback only echoes reasoning if (data?.id) (openAIEndpoint.ts#L160). The id comes only from cot_id, reasoning_opaque or signature (thinkingUtils.ts#L30). Standard OpenAI-compatible servers (SGLang, vLLM and similar) stream reasoning_content with no id, so every later request carries the assistant message's content and tool_calls, but no reasoning.
Models that expect their reasoning back on every assistant message degrade badly. With Kimi K3 served by SGLang, the model reasoned for 68 tokens on the first step and exactly 3 tokens on every later step. After six steps it repeated the same turn until cancelled. Replaying the second request with the first step's reasoning added back gave 175, 303 and 244 reasoning tokens, versus 3, 3 and 3 without it. Adding a cot_id to each reasoning delta in a proxy fixes it.
Affected version
1.138.0 and 1.139.1. The code is unchanged on main.
Steps to reproduce the behavior
- Serve a reasoning model behind an OpenAI-compatible server that streams
reasoning_content, for example SGLang with --reasoning-parser, and log the requests it receives
- Add it under
github.copilot.chat.customOAIModels with "thinking": true and "toolCalling": true
- Run an agent request that makes a few tool calls
- The responses stream
reasoning_content, but the assistant messages in later requests have no reasoning_content
Expected behavior
With "thinking": true, reasoning received as reasoning_content is sent back on its assistant message, whether or not the provider supplied an id.
Describe the bug
With a custom Chat Completions model (
github.copilot.chat.customOAIModels,"thinking": true), VS Code parses streamedreasoning_contentbut never sends it back on later requests.OpenAIEndpoint.getCompletionsCallbackonly echoes reasoningif (data?.id)(openAIEndpoint.ts#L160). The id comes only fromcot_id,reasoning_opaqueorsignature(thinkingUtils.ts#L30). Standard OpenAI-compatible servers (SGLang, vLLM and similar) streamreasoning_contentwith no id, so every later request carries the assistant message'scontentandtool_calls, but no reasoning.Models that expect their reasoning back on every assistant message degrade badly. With Kimi K3 served by SGLang, the model reasoned for 68 tokens on the first step and exactly 3 tokens on every later step. After six steps it repeated the same turn until cancelled. Replaying the second request with the first step's reasoning added back gave 175, 303 and 244 reasoning tokens, versus 3, 3 and 3 without it. Adding a
cot_idto each reasoning delta in a proxy fixes it.Affected version
1.138.0 and 1.139.1. The code is unchanged on
main.Steps to reproduce the behavior
reasoning_content, for example SGLang with--reasoning-parser, and log the requests it receivesgithub.copilot.chat.customOAIModelswith"thinking": trueand"toolCalling": truereasoning_content, but the assistant messages in later requests have noreasoning_contentExpected behavior
With
"thinking": true, reasoning received asreasoning_contentis sent back on its assistant message, whether or not the provider supplied an id.