Post-Hoc Recovery Evaluation: Measuring Agent Recovery After Unsafe Tool Execution
A research project that makes recovery a first-class evaluation object: injecting controlled, sandboxed unsafe-execution events into agent benchmarks and measuring...
Filter ideas by difficulty level — beginner to advanced.
Showing 41 ideas in intermediate
A research project that makes recovery a first-class evaluation object: injecting controlled, sandboxed unsafe-execution events into agent benchmarks and measuring...
A research project that audits the measurement instruments themselves: applying the ABC validity-checklist methodology to agent-security benchmarks to find task-validity...
A research project that audits the auditors: measuring how position, verbosity, and self-preference biases distort LLM-judge judgments of safety violations...
A research project that treats the decision to execute a tool call as a learnable risk-aware policy — training context-conditional...
A research project that measures, for the first time at registry scale, how much tool-poisoning risk actually exists across public...
A research project that measures the full trade-off curve between prompt-injection resistance and task utility for LLM agent defenses —...
Build a local-first transcription workbench that turns speech into text entirely on-device — open Whisper-class models, measurable evaluation against the...
A team-oriented product that batch-validates document libraries against PDF/UA and PDF/A rules using the open-source veraPDF engine — prioritized reports,...
Build a degradation-prediction system that learns from multivariate industrial sensor time series and issues maintenance-support alerts — a supervised remaining-useful-life...