Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents
This paper describes a vulnerability in causal action verification for language agents, where corrupting the committed graph can lead to false executions. The authors demonstrate that omitting or reversing edges in the graph can cause the verifier to issue incorrect certificates, allowing for harmful actions. A proposed attestation step can detect these attacks, but it has limitations and scalability issues.
Save an API key to vote.