Post-Hoc Recovery Evaluation: Measuring Agent Recovery After Unsafe Tool Execution
A research project that makes recovery a first-class evaluation object: injecting controlled, sandboxed unsafe-execution events into agent benchmarks and measuring...
An AI-powered study tool that helps CS students understand concepts through Socratic questioning and adaptive practice.

Project Idea · Intermediate · Python, LLM
CS students struggle with understanding concepts, not writing code. They can follow tutorials but cannot solve problems independently. Traditional study methods are inefficient.
Build an AI study companion using the Socratic method. Instead of giving answers, it asks questions that guide students toward understanding. It adapts to student level and generates practice problems.
Python FastAPI + React. Socratic questioning engine using GPT-4. Cover data structures, algorithms, recursion.
Adaptive difficulty, spaced repetition, concept maps, 10 CS topics.
In-browser code execution, progress dashboard, student accounts.
ChatGPT gives answers. This tool deliberately does not — it asks questions that guide you to the answer yourself. It also tracks progress and schedules review.
A research project that makes recovery a first-class evaluation object: injecting controlled, sandboxed unsafe-execution events into agent benchmarks and measuring...
A research project that answers the question every failed agent run raises: which step broke it? Building an attributed corpus...
A research project that audits the measurement instruments themselves: applying the ABC validity-checklist methodology to agent-security benchmarks to find task-validity...
Published on September 2, 2026
A team of developers, researchers, and innovators who review and publish practical ideas for builders and creators.
Published on September 2, 2026
A team of developers, researchers, and innovators who review and publish practical ideas for builders and creators.