@hue-run/sdk 0.12.1 and hue-run 0.6.4 or later; the agent itself needs no Hue SDK. For the TypeScript callback path and what a world is, see simulations; for the command reference, see hue eval.
How a mirror works
Each app the agent uses has a mirror at a stable Hue URL:https://app.hue.run/api/sim/gmail.googleapis.com/gmail/v1, and a Gmail MCP mirror is https://app.hue.run/api/sim/gmailmcp.googleapis.com/mcp/v1. The agent calls the mirror exactly as it calls the real app, with a per-run world token in place of the real credential. Hue answers from that run’s private world in the real app’s format, records every call in the world’s ledger, then seals the world when the case ends and scores it.
The mirrors serve the real app’s captured catalog, or the tool list the trace the case was built from recorded, so the agent sees what production sees. Hue adds no tool of its own. An eval-ready agent therefore never receives Hue-native tools: not the deprecated context.tools or context.mcp of a world without a handoff, not the Hue MCP server, not a tool added for the evaluation.
The one-time change
Read each app’s URL and credential through a helper that detects an evaluation and falls back to the production values otherwise.hue eval sets HUE_WORLD_TOKEN only in the process it starts, so its presence means an evaluation. During one, the helper reads the mirror URL from the app’s HUE_SIM_<SURFACE>_URL variable and uses HUE_WORLD_TOKEN as the bearer. It fails closed: a missing variable for an app the agent needs throws before any request, and never falls back to the real app inside an evaluation. It also refuses a case without a world: hue eval --command always sets HUE_EXECUTION_ID, so that variable without HUE_WORLD_TOKEN throws instead of running the agent against the real app. Without either variable it returns the production URL and credential unchanged, so production behavior is identical.
Commit the helper and its call sites. Nothing else changes: not the prompts, the model, the tools or the dependencies.
- TypeScript
- Python
hue-eval.ts
Variable names
Each mirror’s variable isHUE_SIM_<SURFACE>_URL, where the surface is the app’s provider surface id in upper case with every other character replaced by _: Gmail’s MCP mirror is HUE_SIM_GOOGLE_GMAIL_MCP_URL and its REST mirror is HUE_SIM_GOOGLE_GMAIL_REST_URL. HUE_MCP_CONFIG names an owner-only mcpServers file for every MCP mirror of the world, keyed by Hue’s provider instance (such as gmail-primary) rather than by your server names, with HUE_WORLD_TOKEN as the bearer; it fits as-is only when those keys are the names your production configuration uses, since a framework derives tool names from them. HUE_MCP_URL, HUE_MCP_TOKEN and HUE_MCP_EXPIRES_AT carry the first MCP mirror under the names older adapters read. hue eval prints world created per case, not the variable names; to learn a case’s exact names, have the agent log the HUE_SIM_ variable names, never their values, to stderr once.
Call sites
Point each app client at the helper’s result. First find how the codebase reaches each app: anmcpServers file, an MCP client in code, a hosted connector, a REST SDK with a base URL option, or raw HTTP calls. Each is a call site. The snippets use a Gmail MCP mirror and a placeholder readGoogleAccessToken() for the production credential; replace the variable and the production values with your app’s.
An MCP configuration file
For a framework that loads anmcpServers file, keep your production file and its server names, and let the helper write the evaluation copy with each server’s URL and credential replaced by its mirror’s. Tool names stay what production has, a server without a mirror or with a client-specific URL field (such as serverUrl or httpUrl) fails closed, and the copy is removed when the agent exits (dispose() in TypeScript, the end of the with block in Python). Claude Code reads the result with --mcp-config:
claude-code.sh
HUE_MCP_CONFIG has the same common shape, one Streamable HTTP server per provider instance; pass it directly only when its keys are your server names:
mcp.json
OpenAI Agents SDK
- TypeScript
- Python
openai-agents.ts
Vercel AI SDK
vercel-ai-sdk.ts
LangChain
langchain_agent.py
Hosted MCP connectors
When the model provider connects to the MCP server itself, pass it the mirror URL and the world token. The mirror is reachable from the provider’s servers, and the token reaches only this run’s world.- OpenAI Responses API
- Anthropic MCP connector
openai-responses.ts
Direct REST clients
A REST mirror takes the same paths as the real API under its base URL. Give an SDK the mirror as its base URL option, or build requests on it:direct-rest.ts
Date-relative tasks
A world’s clock is the recorded start of the trace it was built from, not the wall clock, so “today”, “tomorrow” ornewer_than:7d should be computed from the world’s date. @hue-run/sdk 0.13.0 and hue-run 0.7.0 pass that date to the agent as HUE_WORLD_NOW, an RFC 3339 timestamp, when Hue sends it with the world. Read it when present and fall back to the wall clock, so the agent is ready for it without a further change.
Run it
hue eval starts your agent’s own command once per case, in a fresh world, with the world’s variables set. Nothing is edited per run. The command contract is small: read {"inputs","config"} as JSON from stdin, run the unchanged production agent once, print the answer on stdout and exit. When your production entry point is a server or takes another input shape, add a thin entry point and commit it with the helper:
agent.js
zod peer and run with a Read and write key that hue login stored in an ignored env file:
HUE_WORLD_ID, HUE_WORLD_TOKEN, one HUE_SIM_<SURFACE>_URL per mirror, HUE_MCP_CONFIG, and HUE_MCP_URL, HUE_MCP_TOKEN and HUE_MCP_EXPIRES_AT for the first MCP mirror, plus HUE_EXECUTION_ID, HUE_CASE_ID and HUE_CASE_KEY. It reads {"inputs","config"} as JSON on stdin and prints its answer on stdout, which Hue stores with the credentials hue eval handed it redacted; keep logs on stderr. HUE_API_KEY and other Hue credentials are removed from the command’s environment, and only the current case’s world sets world variables. Exit code 0 means every case passed; 1 means a case failed, errored or a check was still pending; 2 is a usage error. Rerun with --baseline <previous experimentId> to compare.
To let the Run button in Hue and the launch_local_run MCP tool use the same agent, start a worker instead. It claims runs Hue queues and executes them on your machine with the same handoff:
hue eval’s own process receives the same handoff as context.world: surfaces, token, env and mcpConfig. Pass those to the agent the same way, or spawn it with agentEnvironment and writeMcpConfig; do not hand it context.tools or context.mcp. Python drives worlds with the hue_sdk.environment helpers agent_environment and mcp_config_file. A world the simulation gateway does not serve has no handoff and sets no HUE_WORLD_TOKEN; it offers only the deprecated Hue-native capability, which an eval-ready agent does not use. The helper refuses such a case, because HUE_EXECUTION_ID is set without HUE_WORLD_TOKEN, rather than running the agent against the real app. See worlds with a handoff.
Verify before you finish
- Production unchanged. With
HUE_WORLD_TOKENunset, the helper returns the production URL and credential for every call site, and the application’s existing tests pass. - Fail closed. With
HUE_WORLD_TOKENset and oneHUE_SIM_<SURFACE>_URLunset, and withHUE_EXECUTION_IDset butHUE_WORLD_TOKENunset, the helper throws before any request is sent. - Same tools. The agent’s tool list is the one production has. Nothing was added or removed, and nothing Hue-native was handed to the agent.
- Clean diff. The change is the helper plus its call sites, and no Hue variable or value is in production configuration, secrets or a committed file.
- A real run. You have the run URL, the printed verdicts, the world’s recorded calls on the cases that needed them, and stdout that carried no credential.
Production safety
- The world variables exist only in the process
hue evalstarts, andhue evalremoves theHUE_MCP_CONFIGfile after each case. Production never setsHUE_WORLD_TOKENorHUE_EXECUTION_ID, so the helper takes the production branch there and reads your real credential exactly as before. - Never copy
HUE_WORLD_TOKEN,HUE_SIM_*orHUE_MCP_*values into production secrets or a committed file. They belong to one case and are useless and misleading anywhere else. - A world token reaches only one simulated world in one project and expires with it. It is never a credential for the real app, and the real app never sees it.
- Keep the project key out of the agent:
hue evalstripsHUE_API_KEYfrom the command’s environment by default. Pass--allow-hue-credentialsonly for an agent that must call Hue’s own API.
If the agent hard-codes real URLs
An agent that keeps a real URL calls the real app during the evaluation, outside the world. Hue records none of those calls, so the world seals with nothing to score, while the real app receives the writes. Make every app URL and credential configurable through the helper above. The world’sno_calls flag is advisory: an agent can also stop before its first request or invent an answer without using an app. On a case that needed app calls, a no_calls world together with an answer that claims the agent acted has two possible causes, a call site that still points at production or an agent that made no request. Check where the agent’s requests went, from its own logs, before changing a call site.