docs(adr-002): re-adjudicate Groq model to servable qwen/qwen3.8-27b (APA-47) - #114
Merged
Merged
Conversation
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…(APA-47)
ADR-002's Option A adjudication contracted llama-3.1-8b-instant, matching
the code default. That model has since been retired from Groq, so the
documented production config would fail every inference call with HTTP 404
model_not_found. Re-adjudicate to a currently servable model.
Verified live against the configured provider (defaultGroqBaseURL) using
the real key from the untracked .env, replicating the client's exact wire
format (max_tokens, temperature 0, stream false):
- qwen/qwen3.8-27b -> HTTP 200, echoed model qwen/qwen3.8-27b,
non-empty choices[0].message.content
- llama-3.1-8b-instant -> HTTP 404 model_not_found (retired)
- account model list: 11 models, zero llama-3.x, qwen/qwen3.8-27b present
Config and docs only. No ModelClient behavior change: no response_format
or JSON mode, no reasoning parameter, no prompt change, no model-specific
response parsing, no new dependencies. The model remains untrusted input
behind the unchanged deterministic validation/grounding boundary.
Moves the canonical model string in the four places that must agree: the
groq_model.go default, the README provider line, the ADR-002 ## Decision
line, and .env.example GROQ_MODEL (which had drifted to the stale
pre-adjudication openai/gpt-oss-20b value).
The two APA-45 test literals that pin the canonical string are renamed to
the new value. Renames only; no assertion weakened or deleted. RED was
confirmed first: flipping the literals alone failed 5 tests across both
packages, proving the coupling assertions are load-bearing.
Gates: gofmt clean, go vet clean, go build clean, go test ./... 43 pkgs
0 fail (6 pre-existing self-skips: TestOutbox_* x3, TestLaunchLive_* x3),
uv run pytest tests/unit 27 passed, ruff check clean, and
git diff main -- eval/ tests/ empty.
Aparnap2
force-pushed
the
docs/adr-002-qwen-re-adjudication
branch
from
September 28, 2026 22:28
c0fea27 to
ad0d631
Compare
Aparnap2
marked this pull request as ready for review
September 28, 2026 22:28
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Re-adjudicate ADR-002 to a currently servable production model. Config + docs only.
ADR-002's Option A adjudication contracted
llama-3.1-8b-instant, matching the code default. That model has since been retired from Groq, so the documented production config would fail every inference call. Production model is nowqwen/qwen3.8-27b.Canonical string moved in the four places that must agree:
apps/api/internal/investigate/orchestrate/groq_model.gollama-3.1-8b-instantqwen/qwen3.8-27bREADME.mdllama-3.1-8b-instantqwen/qwen3.8-27bdocs/adr/002-groq-llm-provider.md(## Decision)llama-3.1-8b-instantqwen/qwen3.8-27b.env.exampleopenai/gpt-oss-20bqwen/qwen3.8-27b.env.examplehad additionally drifted to the stale pre-adjudication value; it now matches code/ADR/README.No ModelClient behavior change. No
response_format/JSON mode, no reasoning parameter, no prompt change, no model-specific response parsing, no new dependencies.go.mod/go.sumuntouched. The model remains untrusted input behind the unchanged deterministic validation/grounding boundary. No changes toeval/, eval-v1, the A–P corpus, E0–E3, hard gates, scoring, or any APA-11/12, retry, tenant, OCR, workflow, or agent contract.Evidence
Servability confirmed before any edit, through the actual configured provider (
defaultGroqBaseURL) with the real key from the untracked.env, replicating the client's exact wire format (max_tokens,temperature: 0,stream: false):choices[0].message.contentqwen/qwen3.8-27bqwen/qwen3.8-27b"OK")llama-3.1-8b-instantmodel_not_foundAccount model list: 11 models, zero
llama-3.x,qwen/qwen3.8-27bpresent. Premise independently corroborated.The
.envkey was never printed, never staged, never committed; it is still gitignored and untracked.Test expectation updates (renames only, no weakening). Two APA-45 literals that pin the canonical string were renamed to the new value:
groq_env_contract_test.go→canonicalModelStringgroq_provider_selection_test.go→canonicalGroqModelProvenance comments updated to record the APA-47 re-adjudication. No assertion weakened, skipped, or deleted. RED was confirmed first: flipping only the literals failed 5 tests across both packages (
TestGroqContract_DefaultModelIsCanonicalString,TestGroqContract_DocsAgreeOnCanonicalModelString,TestGroqContract_EmptyModelResolvesToCanonicalDefault,TestAgentProvider_GroqEnvSelectsGroqClientNotMock,TestAgentProvider_OutboundCarriesCanonicalModelString) — proving the coupling assertions are load-bearing. All pass after the move, including the ADR/README/code coupling testTestGroqContract_DocsAgreeOnCanonicalModelString.overrideGroqModel = "llama-3.3-70b-versatile"was deliberately not changed: it is a deliberately non-canonical override fixture on a loopback stub that provesGROQ_MODELresolution wins, never dialled.Blockers
None for this change. Two out-of-scope stale references found, not touched (they are outside this file set and one is a dated historical record — flagging rather than silently rewriting):
docs/specs/09_DECISION_LOG.md:223,226— records the 2026-09-26 Option A canonical string. Wants a new dated APA-47 entry, not a rewrite of the old one.09is architect-owned.docs/specs/10_PHASE3_E2E_QUALIFICATION.md:161— names the old default. Note this line also says "retry 3x", which was already stale before this PR (the client makes 2 attempts total); pre-existing, unrelated.Next step
Review and merge. Suggested follow-up: a dated APA-47 entry in
09_DECISION_LOG.md(architect-owned) so the decision log does not contradict the ADR.Gates on
c0fea27:gofmt -lclean ·go vet ./...clean ·go build ./...clean ·go test ./...43 pkgs, 0 fail ·uv run pytest tests/unit27 passed ·uv run ruff check .clean ·git diff main -- eval/ tests/empty.