Former sec researcher here. A LOT of the failures right now aren't about an agent doing anything wrong. More like agent that's getting to seeing too much
mail -> search -> docs -> agent
3 hops, now the agent has vaccumed up context that no human could have gathered manually, all in seconds hooked right into the Graph (yes, copilot, talking about u)
Delegation depth for sure, next add context accumulation because it deserves its own ctageory for study
Feels like one of those few benchmarks measuring smth fundamentally different vs adding more prompt injection samples. Will share around
Former sec researcher here. A LOT of the failures right now aren't about an agent doing anything wrong. More like agent that's getting to seeing too much
3 hops, now the agent has vaccumed up context that no human could have gathered manually, all in seconds hooked right into the Graph (yes, copilot, talking about u)
Delegation depth for sure, next add context accumulation because it deserves its own ctageory for study
Feels like one of those few benchmarks measuring smth fundamentally different vs adding more prompt injection samples. Will share around