HoosierHelp
HoosierHelp is an interactive benchmark for evaluating LLM agents in social service navigation. Agents interact with simulated users, issue structured resource-search calls, and select final resources from 3,971 Indiana public social service resources. The evaluation focuses on constraint grounding and handling non-ideal user interactions.
- Released
- 2026-07-03
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Existing benchmarks do not capture the interaction complexity and constraint-grounding demands of social service navigation. HoosierHelp addresses this gap by simulating realistic user behaviors, providing a basis for assessing agent reliability in a high-stakes domain where incorrect resource recommendations can have serious consequences.
Motivation
Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.