FreshCache-Bench
FreshCache-Bench provides 8,072 base queries across five freshness classes with ground truth staleness labels from web snapshots at 1, 12, 24 hours, and 7 days, expanded to 31,201 queries via paraphrase generation, for evaluating semantic caching in retrieval-augmented LLMs.
- Released
- 2026-07-05
- Readiness
- Paper only
- Primary field
- General AI
Why it matters
Semantic caching for RAG lacks standardized evaluation of freshness; this benchmark addresses the gap by quantifying stale error and search API savings, enabling comparison of caching strategies in terms of cost and correctness.
Motivation
Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar queries, but most existing methods do not model the time-varying freshness of open-web evidence.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.