Benchmark Radar
AI BENCHMARK PROFILE

HoosierHelp

General AIRobotics & Embodied Intelligence

HoosierHelp is an interactive benchmark for evaluating LLM agents in social service navigation. Agents interact with simulated users, issue structured resource-search calls, and select final resources from 3,971 Indiana public social service resources. The evaluation focuses on constraint grounding and handling non-ideal user interactions.

Released
2026-07-03
Readiness
Paper only
Primary field
General AI

Why it matters

Existing benchmarks do not capture the interaction complexity and constraint-grounding demands of social service navigation. HoosierHelp addresses this gap by simulating realistic user behaviors, providing a basis for assessing agent reliability in a high-stakes domain where incorrect resource recommendations can have serious consequences.

Motivation

Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.