StreamProfileBench
StreamProfileBench evaluates LLMs on fine-grained streaming user profiling. Models maintain a rolling persona summary from a stream of user posts and predict which tags from a candidate pool the user will engage with next. Includes over 120,000 posts from 7,000+ users across five Chinese platforms with metrics for recall, novelty, stability, and error rates.
- Released
- 2026-05-25
- Readiness
- Runnable
- Primary field
- General AI
Why it matters
Existing user profiling benchmarks use static data, failing to capture real-world streaming UGC and rapidly evolving interests. StreamProfileBench provides a dynamic evaluation that measures plasticity-stability balance, revealing conservative bias in LLMs and enabling practical improvements for personalized systems.
Motivation
Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots.
Primary resources
Benchmark Radar records only publicly supported details and links back to primary sources for verification.