Research1 min read
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We in...
From arXiv cs.AI