Skip to content

feat: DeepSeek as a supported provider - #22

Merged
babaliauskas merged 15 commits into
mainfrom
feat/deepseek-provider
Sep 30, 2026
Merged

babaliauskas merged 15 commits into
mainfrom
feat/deepseek-provider

Conversation

@babaliauskas

@babaliauskas babaliauskas commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Adds DeepSeek (deepseek-flash, deepseek-v4-pro) as a first-class provider.

Draft because the live check against the DeepSeek API hasn't run yet (no key was available). The checklist is below.

What changes

  • Registry.
    • New provider deepseek, with DEEPSEEK_API_KEY checked before run/compare and shown as a row in doctor.
    • Bare deepseek-* ids resolve to deepseek/…. This is what a capture records when the app calls api.deepseek.com through the OpenAI client; before, LiteLLM couldn't route these ids.
    • The judge-family warning now works for a DeepSeek judge grading a DeepSeek arm.
    • DeepSeek served by another host (azure_ai/, hosted_vllm/, openrouter/) stays other, so it isn't asked for DEEPSEEK_API_KEY.
  • Tool calls. Any id containing deepseek is parsed as OpenAI-shaped. Previously every DeepSeek tool eval raised ModelError.
  • Pricing. Recorded calls are priced under the provider-prefixed key first. Bare deepseek-flash was priced at $0 before. Side effect: a few niche Gemini keys change price (gemini-exp-1206 now prices at $0, and image-model cache/tier prices differ); text-model prices are unchanged.
  • Thinking mode (new models/deepseek.py). Both models think by default. Thinking mode silently ignores temperature, and it rejects tool requests with a 400 unless earlier assistant turns carry reasoning_content.
    • DeepSeek arms, and a DeepSeek judge, now get the report's non-determinism banner.
    • Replayed assistant turns (tool rounds and chat history) get the single-space placeholder DeepSeek accepts.
    • EvalShift never switches thinking off.
  • init --provider deepseek. Scaffolds deepseek-flash as source and deepseek-v4-pro as target and judge. The semantic block ships commented out because DeepSeek has no embedding endpoint.
  • Docs. DOCS.md, llms-full.txt, docs/ and the CHANGELOG are updated, including a FAQ entry "Does EvalShift work with DeepSeek?". New docs-currency tests make the docs name every registry key env var and every init --provider choice.

Plan: docs/superpowers/plans/2026-09-30-deepseek-provider.md. Each task had a spec and quality review, and the whole branch had a final review.

CI note

Rebased onto main after #21 (the SQLAlchemy 2.1 cache fix) merged. CI is green on Python 3.11–3.14.

Live verification checklist (needs DEEPSEEK_API_KEY; the judge run also needs GEMINI_API_KEY)

  • Smoke script. Run DEEPSEEK_API_KEY=... uv run python scripts/smoke_live_tools.py. For each DeepSeek model, expect:
    • tool calls and a cost above $0 for the tool prompts;
    • round 1: with no 400, which proves the backfill;
    • history: text=True, which shows the placeholder is harmless on a request without tools.
    • Commit the new tests/unit/fixtures/tool_responses/deepseek/*_live.json files.
  • Backfill proof.
    • Temporarily disable the if "messages" in kwargs and thinking_by_default(canonical): block in src/evalshift_cli/models/client.py (around lines 739–744).
    • Rerun the script for deepseek-flash only. round 1 should now FAIL with DeepSeek's "reasoning_content … must be passed back" 400.
    • Restore the block. If round 1 passes without it, DeepSeek has relaxed the rule; keep the block and note it here.
  • test-call. Run DEEPSEEK_API_KEY=... uv run evalshift test-call --model deepseek-flash --max-tokens 2048 and expect a reply and a non-zero cost.
    • Also try the default --max-tokens 256. If it returns an empty, thinking-only reply, reasoning tokens count toward max_tokens, and the FAQ should say so.
  • Compare. Run cd examples/agent && DEEPSEEK_API_KEY=... GEMINI_API_KEY=... uv run evalshift compare --from deepseek-flash --to deepseek-v4-pro --yes. Expect:
    • the report is written;
    • the banner names both DeepSeek arms;
    • tool-call scores are populated;
    • evalshift doctor shows DEEPSEEK_API_KEY set.
  • DeepSeek judge. Copy examples/agent to a scratch directory, set judge_model: deepseek-v4-pro on an llm_judge entry, and rerun the compare. Expect:
    • the verdicts are populated;
    • the banner lists deepseek/deepseek-v4-pro exactly once.

Release follows the plan's Phase E: CLI 1.2.0, then the action pin PR, then the sibling PRs and the site npm run sync:llms.

🤖 Generated with Claude Code

babaliauskas and others added 15 commits September 30, 2026 12:10
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Bare deepseek-* ids now resolve to deepseek/…, DEEPSEEK_API_KEY is pre-checked and shown by doctor, and a DeepSeek judge on a DeepSeek arm is flagged as one family.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Thinking mode ignores temperature, so DeepSeek arms now get the non-determinism banner; replayed assistant turns get the placeholder reasoning_content DeepSeek requires on tools requests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Also print the required provider API key env var in init's next-steps
output unconditionally (previously only shown under --ci), so a plain
'init --provider deepseek' tells the user which key to export.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Controller ruling: the plan never asked for a user-visible change to
init's next-steps for every provider (it would need docs). Revert the
unconditional API-key line added in 1769343 and drop the now-unsupported
stdout assertion from the DeepSeek test; the key mapping stays pinned by
test_deepseek_workflow_uses_the_deepseek_key (--ci workflow) and
test_every_init_provider_key_is_the_registry_key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…stry

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ke script

Steps 2-4 (live runs with a real DEEPSEEK_API_KEY, the with/without-backfill 400 check, and the end-to-end CLI run) are deferred to the maintainer per controller ruling R1: no provider key is available in this environment.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pin that a DeepSeek thinking arm lands on the non-determinism list through
the real honors_temperature path, and that a text-only chat-history
assistant turn is sent with the placeholder reasoning_content without
mutating the caller's list.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The reasoning_content backfill runs on every assistant turn replayed to a
DeepSeek thinking model, chat history as well as tool rounds; say so in the
client docstrings and comment instead of "verbatim" / "unconditional", and
say why (required on tool requests, ignored otherwise). List DeepSeek as the
second exception in the capabilities module docstring, stop claiming
honors_temperature cannot drift from unsupported_params, and note that
thinking_by_default over-reports for opt-in-thinking models.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A DeepSeek judge accepts temperature and ignores it, so the client never
sees a rejection and the judge never reached the non-determinism banner.
Check each configured judge with honors_temperature, as the run's arms
are, and merge its canonical id alongside the runtime rejections.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The FAQ said the placeholder reasoning chain rides only on tool requests;
it rides on every assistant turn replayed from the recording, chat history
included. Say why (required with tools, ignored without), note that a
DeepSeek judge is marked non-deterministic too, and tell Ollama users to
prefix a deepseek-r1 arm with ollama/. Lead the OpenAI-compatible server
list in the SDK mirrors with DeepSeek, as the SDK docs now do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The DeepSeek backfill also rides on plain chat-history assistant turns,
which the live script never exercised. Add a user/assistant/user
complete_messages check per model after round 1, and bring the module
docstring up to date with the round-1 and history checks and the real
<prompt_name>_live.json fixture path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Record the opt-in-thinking caveat on decision 4, replace decision 5's
"changes nothing" with the measured pricing effect (ruling R10), fix the
test-call invocation in Task A7, and add the text-only history and
DeepSeek-judge end-to-end checks to A7.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@babaliauskas
babaliauskas marked this pull request as ready for review September 30, 2026 10:13
@babaliauskas
babaliauskas merged commit 6928943 into main Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant