Is your feature request related to a problem?
The v2 self-driving eval currently lacks a UI for starting or monitoring the prompt iteration loop, making it an API-only feature. This limits user accessibility and understanding of the process.
Describe the solution you'd like
- Add a "Run iteration loop" entry point:
dataset_id, experiment_name, config_id, config_version, max_rounds
- Handle the
202 response with EvaluationIterationRunImmediatePublic details (iteration_run_id, status)
- Implement a webhook receiver for the
EvaluationIterationReportPublic callback, aligned with the prompt-improvement pattern
- Display round-by-round history, the best round, and terminal
stop_reason values (ceiling_reached, max_rounds_reached, round_failed)
Original issue
Is your feature request related to a problem?
v2 ships a self-driving eval → improve-prompt → eval loop (POST /api/v2/evaluations/iterations). There is no way to start one or watch it from the console, so the feature is API-only today.
Describe the solution you'd like
- Add a "Run iteration loop" entry point:
dataset_id, experiment_name, config_id, config_version, max_rounds
- Handle the
202 response and its EvaluationIterationRunImmediatePublic handle (iteration_run_id, status)
- Add a webhook receiver for the
EvaluationIterationReportPublic callback, following the prompt-improvement pattern
- Show round-by-round history, the best round, and the terminal
stop_reason (ceiling_reached, max_rounds_reached, round_failed)
Comments: Stop-score is the mean of Adherence to Ground Truth and Adherence to Prompt; Knowledge Base is recorded per round for visibility only — worth reflecting in the UI so the numbers aren't confusing.
Parent: #265
Is your feature request related to a problem?
The v2 self-driving eval currently lacks a UI for starting or monitoring the prompt iteration loop, making it an API-only feature. This limits user accessibility and understanding of the process.
Describe the solution you'd like
dataset_id,experiment_name,config_id,config_version,max_rounds202response withEvaluationIterationRunImmediatePublicdetails (iteration_run_id,status)EvaluationIterationReportPubliccallback, aligned with the prompt-improvement patternstop_reasonvalues (ceiling_reached,max_rounds_reached,round_failed)Original issue
Is your feature request related to a problem?
v2 ships a self-driving eval → improve-prompt → eval loop (
POST /api/v2/evaluations/iterations). There is no way to start one or watch it from the console, so the feature is API-only today.Describe the solution you'd like
dataset_id,experiment_name,config_id,config_version,max_rounds202response and itsEvaluationIterationRunImmediatePublichandle (iteration_run_id,status)EvaluationIterationReportPubliccallback, following the prompt-improvement patternstop_reason(ceiling_reached,max_rounds_reached,round_failed)Comments: Stop-score is the mean of Adherence to Ground Truth and Adherence to Prompt; Knowledge Base is recorded per round for visibility only — worth reflecting in the UI so the numbers aren't confusing.
Parent: #265