Post-Hoc Recovery Evaluation: Measuring Agent Recovery After Unsafe Tool Execution
A research project that makes recovery a first-class evaluation object: injecting controlled, sandboxed unsafe-execution events into agent benchmarks and measuring...
Explore project ideas, product concepts, startup opportunities, and technology innovations — all in one place.
A research project that makes recovery a first-class evaluation object: injecting controlled, sandboxed unsafe-execution events into agent benchmarks and measuring...
A research project that answers the question every failed agent run raises: which step broke it? Building an attributed corpus...
A research project that audits the measurement instruments themselves: applying the ABC validity-checklist methodology to agent-security benchmarks to find task-validity...
A research project that audits the auditors: measuring how position, verbosity, and self-preference biases distort LLM-judge judgments of safety violations...
A research project that treats the decision to execute a tool call as a learnable risk-aware policy — training context-conditional...
A research project that studies how prompt injections persist across agent sessions — building a stateful evaluation harness that measures...
Find projects in your area of interest.
Find projects built with a specific technology.
Browse practical project, product, and business ideas organized by domain and technology.
Find ideas by type, domain, technology, or difficulty level to match your skills and interests.
Each idea includes a clear problem statement, proposed solution, and key features to guide implementation.
Browse our growing collection of practical ideas for developers, entrepreneurs, and students.