Too Good To Be True
A hard ai_coding interview practice problem on DataDriven. Write and execute real ai_coding code with instant grading.
- Domain
- ai_coding
- Difficulty
- hard
- Seniority
- L6
Problem
Last week's RAG eval reported recall@5 = 0.97 (up from 0.62) with no documented index or model change. Finance wants to expand the agent rollout to all enterprise tenants based on this number. The data team is suspicious. Reproduce by running tests/test_eval_sanity.py. Fix what the test surfaces and whatever else you find while tracing the data flow. Be ready to defend why the new recall number is trustworthy when the interviewer asks 'what would you check next?'
Summary
Suspiciously flawless.
Practice This Problem
Solve this ai_coding problem with real code execution. DataDriven runs your solution and grades it automatically.