feat: DeepSeek as a supported provider - #22
Merged
Merged
Conversation
This was referenced Sep 30, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bare deepseek-* ids now resolve to deepseek/…, DEEPSEEK_API_KEY is pre-checked and shown by doctor, and a DeepSeek judge on a DeepSeek arm is flagged as one family. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Thinking mode ignores temperature, so DeepSeek arms now get the non-determinism banner; replayed assistant turns get the placeholder reasoning_content DeepSeek requires on tools requests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Also print the required provider API key env var in init's next-steps output unconditionally (previously only shown under --ci), so a plain 'init --provider deepseek' tells the user which key to export. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Controller ruling: the plan never asked for a user-visible change to init's next-steps for every provider (it would need docs). Revert the unconditional API-key line added in 1769343 and drop the now-unsupported stdout assertion from the DeepSeek test; the key mapping stays pinned by test_deepseek_workflow_uses_the_deepseek_key (--ci workflow) and test_every_init_provider_key_is_the_registry_key. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…stry Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ke script Steps 2-4 (live runs with a real DEEPSEEK_API_KEY, the with/without-backfill 400 check, and the end-to-end CLI run) are deferred to the maintainer per controller ruling R1: no provider key is available in this environment. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pin that a DeepSeek thinking arm lands on the non-determinism list through the real honors_temperature path, and that a text-only chat-history assistant turn is sent with the placeholder reasoning_content without mutating the caller's list. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The reasoning_content backfill runs on every assistant turn replayed to a DeepSeek thinking model, chat history as well as tool rounds; say so in the client docstrings and comment instead of "verbatim" / "unconditional", and say why (required on tool requests, ignored otherwise). List DeepSeek as the second exception in the capabilities module docstring, stop claiming honors_temperature cannot drift from unsupported_params, and note that thinking_by_default over-reports for opt-in-thinking models. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A DeepSeek judge accepts temperature and ignores it, so the client never sees a rejection and the judge never reached the non-determinism banner. Check each configured judge with honors_temperature, as the run's arms are, and merge its canonical id alongside the runtime rejections. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The FAQ said the placeholder reasoning chain rides only on tool requests; it rides on every assistant turn replayed from the recording, chat history included. Say why (required with tools, ignored without), note that a DeepSeek judge is marked non-deterministic too, and tell Ollama users to prefix a deepseek-r1 arm with ollama/. Lead the OpenAI-compatible server list in the SDK mirrors with DeepSeek, as the SDK docs now do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The DeepSeek backfill also rides on plain chat-history assistant turns, which the live script never exercised. Add a user/assistant/user complete_messages check per model after round 1, and bring the module docstring up to date with the round-1 and history checks and the real <prompt_name>_live.json fixture path. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Record the opt-in-thinking caveat on decision 4, replace decision 5's "changes nothing" with the measured pricing effect (ruling R10), fix the test-call invocation in Task A7, and add the text-only history and DeepSeek-judge end-to-end checks to A7. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
babaliauskas
force-pushed
the
feat/deepseek-provider
branch
from
September 30, 2026 10:11
facf4b8 to
1d85a2a
Compare
babaliauskas
marked this pull request as ready for review
September 30, 2026 10:13
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds DeepSeek (
deepseek-flash,deepseek-v4-pro) as a first-class provider.Draft because the live check against the DeepSeek API hasn't run yet (no key was available). The checklist is below.
What changes
deepseek, withDEEPSEEK_API_KEYchecked beforerun/compareand shown as a row indoctor.deepseek-*ids resolve todeepseek/…. This is what a capture records when the app callsapi.deepseek.comthrough the OpenAI client; before, LiteLLM couldn't route these ids.azure_ai/,hosted_vllm/,openrouter/) staysother, so it isn't asked forDEEPSEEK_API_KEY.deepseekis parsed as OpenAI-shaped. Previously every DeepSeek tool eval raisedModelError.deepseek-flashwas priced at $0 before. Side effect: a few niche Gemini keys change price (gemini-exp-1206now prices at $0, and image-model cache/tier prices differ); text-model prices are unchanged.models/deepseek.py). Both models think by default. Thinking mode silently ignorestemperature, and it rejects tool requests with a 400 unless earlier assistant turns carryreasoning_content.init --provider deepseek. Scaffoldsdeepseek-flashas source anddeepseek-v4-proas target and judge. The semantic block ships commented out because DeepSeek has no embedding endpoint.docs/and the CHANGELOG are updated, including a FAQ entry "Does EvalShift work with DeepSeek?". New docs-currency tests make the docs name every registry key env var and everyinit --providerchoice.Plan:
docs/superpowers/plans/2026-09-30-deepseek-provider.md. Each task had a spec and quality review, and the whole branch had a final review.CI note
Rebased onto
mainafter #21 (the SQLAlchemy 2.1 cache fix) merged. CI is green on Python 3.11–3.14.Live verification checklist (needs
DEEPSEEK_API_KEY; the judge run also needsGEMINI_API_KEY)DEEPSEEK_API_KEY=... uv run python scripts/smoke_live_tools.py. For each DeepSeek model, expect:round 1:with no 400, which proves the backfill;history: text=True, which shows the placeholder is harmless on a request without tools.tests/unit/fixtures/tool_responses/deepseek/*_live.jsonfiles.if "messages" in kwargs and thinking_by_default(canonical):block insrc/evalshift_cli/models/client.py(around lines 739–744).deepseek-flashonly.round 1should now FAIL with DeepSeek's "reasoning_content … must be passed back" 400.DEEPSEEK_API_KEY=... uv run evalshift test-call --model deepseek-flash --max-tokens 2048and expect a reply and a non-zero cost.--max-tokens 256. If it returns an empty, thinking-only reply, reasoning tokens count towardmax_tokens, and the FAQ should say so.cd examples/agent && DEEPSEEK_API_KEY=... GEMINI_API_KEY=... uv run evalshift compare --from deepseek-flash --to deepseek-v4-pro --yes. Expect:evalshift doctorshowsDEEPSEEK_API_KEY set.examples/agentto a scratch directory, setjudge_model: deepseek-v4-proon anllm_judgeentry, and rerun the compare. Expect:deepseek/deepseek-v4-proexactly once.Release follows the plan's Phase E: CLI 1.2.0, then the action pin PR, then the sibling PRs and the site
npm run sync:llms.🤖 Generated with Claude Code