Benchmark Radar
AI BENCHMARK PROFILE

StreamProfileBench

General AIKnowledge & ReasoningStreamProfileBench Team

StreamProfileBench evaluates LLMs on fine-grained streaming user profiling. Models maintain a rolling persona summary from a stream of user posts and predict which tags from a candidate pool the user will engage with next. Includes over 120,000 posts from 7,000+ users across five Chinese platforms with metrics for recall, novelty, stability, and error rates.

Released
2026-05-25
Readiness
Runnable
Primary field
General AI

Why it matters

Existing user profiling benchmarks use static data, failing to capture real-world streaming UGC and rapidly evolving interests. StreamProfileBench provides a dynamic evaluation that measures plasticity-stability balance, revealing conservative bias in LLMs and enabling practical improvements for personalized systems.

Motivation

Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots.

Primary resources

Benchmark Radar records only publicly supported details and links back to primary sources for verification.