Automated Failure Attribution in Long-Horizon Agent Runs
A research project that answers the question every failed agent run raises: which step broke it? Building an attributed corpus...
Filter ideas by difficulty level — beginner to advanced.
Showing 20 ideas in Advanced
A research project that answers the question every failed agent run raises: which step broke it? Building an attributed corpus...
A research project that studies how prompt injections persist across agent sessions — building a stateful evaluation harness that measures...
A research project that trains unsupervised models on benign agent trajectories to detect multi-step prompt-injection campaigns before the harmful action...
A research project that measures whether prompt-injection defenses keep working when both the attack strategy and the agent architecture change...
A research project that builds a reproducible, execution-scored evaluation harness for desktop AI agents — sandboxed environments, scripted task definitions,...
A research-grade NILM project that estimates appliance-level electricity consumption from a single whole-home power signal — feature extraction, sequence modeling,...
A research-grade chemometrics project that predicts soil properties like organic carbon and pH from diffuse-reflectance spectra using the EU's public...
A research tool for exploring and comparing crop trait and germplasm data from public agricultural databases — normalized schemas, provenance...
Build a defensive tool that inspects Docker image layers, inventories installed packages, matches known vulnerabilities against an updatable feed, and...