Do LLM Safety Judges Need Full Trajectories to Detect Unsafe Behaviour?
Introduction When an LLM is used to judge whether an AI agent’s behaviour violated a safety specification, does the amount of the agent’s trajectory the judge can see affect its ability to detect violations? Specifically, does access to intermediate tool calls and tool results — not just the agent’s final response — improve detection of unsafe behaviour, and can a selective subset of the trajectory recover most of that benefit at lower information cost? ...