LLM-Judge Reliability for Agent Safety Violations on Attacked Trajectories
A research project that audits the auditors: measuring how position, verbosity, and self-preference biases distort LLM-judge judgments of safety violations...
Find ideas built with specific technologies and frameworks.
Showing 8 ideas in llm
A research project that audits the auditors: measuring how position, verbosity, and self-preference biases distort LLM-judge judgments of safety violations...
A research project that treats the decision to execute a tool call as a learnable risk-aware policy — training context-conditional...
Build a flashcard app with a transparent spaced-repetition scheduler and optional LLM-generated explanations of missed cards — scheduling math you...
Build a tool that crawls a site, runs automated accessibility checks against common WCAG-oriented rules, and produces a prioritized report...
Build a system that turns a learner's goals and current knowledge into a step-by-step study path — explainable recommendations, progress...
Build a developer tool that reads your functions, proposes meaningful test cases — happy paths, edge cases, failure modes —...
Build a CLI that analyzes your Git diff and suggests clear Conventional Commits-style messages, with human review always in control.
An AI-powered study tool that helps CS students understand concepts through Socratic questioning and adaptive practice.