Expand evaluation beyond LLM-judged responses.
Potential task types:
- exact match
- contains
- regex
- structured JSON
- file diff
- schema validation
- safe command-result validation
- composite checks
Each type should have a clear schema, useful failure output and tests.
Do not execute arbitrary repository or model-generated code on the host.
Expand evaluation beyond LLM-judged responses.
Potential task types:
Each type should have a clear schema, useful failure output and tests.
Do not execute arbitrary repository or model-generated code on the host.