Why is your model being careful with email bodies but reckless with bank accounts?

Role-stratified per-field conformal risk control calibrates LLM tool calls by semantic argument role rather than certifying each action as one aggregate object. In AgentDojo and InjecAgent, the method assigns separate thresholds and budgets to target, credential, command, selector, control, and content fields, matching certification to where an injection can cause harm.


Read more

Scroll to Top