hue eval runs your own agent against Hue cases and prints the run link and verdicts. Your agent, prompts and model keys stay in your process; Hue never executes the agent. Use this command instead of writing a harness, a runner or tool stand-ins around the SDK.
Before you start
- Node.js 22 or 24 on the machine that runs the command, even for a Python agent.
- A Read and write project key in an ignored env file.
hue login --gitignorewrites.env.hueand ignores it; when the shell has noHUE_API_KEY,hue evalloads./.env.hueby default. Pass--env-file <path>for another file, and add.hue/to.gitignorefor checkpoints. - A start command that reads
{"inputs","config"}as JSON on stdin and prints the answer on stdout. Logs go to stderr. - For cases where the agent calls apps, the one-time mirror helper so its clients reach Hue’s simulated world instead of production.
Run it
Check first, then run the same command without--check:
--check resolves the key’s project, the selection and its evaluators, and whether cases are answer-only or get a world. It creates no run, world or checkpoint. A published case carries its own evaluators. To run an eval set, use --set "<eval set>" and add --scorer-version <id> for each evaluator; Hue’s Run snippet for the set includes them.
- One command, one run. The run holds every selected case: a 100-case eval set is one run with 100 executions. The command prints the
https://app.hue.run/runs/<id>link once the run exists. - Let it finish. Start the command in the background or with a long timeout. If it stops early (Ctrl+C, SIGTERM, a tool timeout or a crash), the run stays open: run the identical command again to resume it from
.hue/eval/<agent-key>. Finished cases are kept, and an interrupted case runs again, in a fresh world for a world case. After a crash, an answer-only case that was mid-run is reported uncertain instead. If it stops while waiting for Hue’s checks, the cases already finished: open the run link instead of rerunning. Never delete that directory to start over. - Exit codes. 0 every case passed; 1 a case failed, errored, is inconclusive or is incomplete; 2 usage error; 130 interrupted.
- Compare. Pass
--baseline <run id|url>to see improvements and regressions against an earlier run.
Answer-only and world cases
The helper fails closed: with an execution but no world token, or a missing mirror URL, a required app connection throws before any request. Outside a Hue eval it returns production values unchanged.
hue eval removes HUE_API_KEY and other Hue credentials from the command’s environment unless you pass --allow-hue-credentials.
Local or deployed
- Local, one run: the one-shot command above.
- Local, launched from Hue:
hue eval --worker --command "<start command>"registers the agent and runs what the Run button orlaunch_local_runstarts, one run per launch, until you stop it. - Deployed: run the worker beside the deployment with a non-production key and a start command that forwards stdin to the deployed agent and prints its answer. Never put world variables or a Read and write key in production configuration.
hue eval for every flag and simulations for how worlds are built and graded.