SocietyBench: Forecasting Counterfactual Social-World Evolution
SocietyBench evaluates LLM agents on forecasting real social-world events by ingesting news/social media timelines and measuring counterfactual event prediction.
Every story tagged with this topic, ordered by date.
SocietyBench evaluates LLM agents on forecasting real social-world events by ingesting news/social media timelines and measuring counterfactual event prediction.
WorldCup Arena prospectively evaluates six frontier LLMs with extended thinking on live 2026 FIFA World Cup forecasting across 104 matches, eliminating memorization risk.
PAST-Bench benchmarks recursive self-improvement in personal AI agents by measuring whether retained preferences, task histories, and skills improve performance over sessions.
Test-time scaling taxonomy clarifies distinct inference algorithms (sampling, voting, search) in reasoning LLMs, standardizing evaluation protocols and compute accounting.
HIVE benchmark reveals voice transcription and keyboard perturbations reduce LLM accuracy across instruction tasks.
HalluTruthQA-4K: 4,000 expert-annotated Arabic QA instances for fine-grained hallucination detection across four knowledge domains.
PRISM meta-workflow converts multivariate time series to multi-channel images for anomaly detection, evaluating vision backbone alternatives.
ANNOTARES dataset for extracting logical structures (conditions and consequences) from German statutory texts.
BanglaWild benchmark evaluates 15 VLMs and OCR systems on 2,535 in-the-wild Bengali scene text images.
MultiGlobeQA: 46K-example multilingual benchmark for geospatial reasoning exposing LLM failures in spatial relations across 14 function families.
ContinualSkillBench evaluates whether LLM agents can continually acquire and reuse skills across 500 interconnected task domains.
ADMITBench: Safety-governed evaluation framework for industrial LLM recommendations assessing evidence support, authority, and plant-specific consequences.
SciRet: Empirical study of hybrid BM25+dense retrieval, reranking, and answer generation for scientific RAG across corpus scales on CORD-19.
GDPevo: benchmark for evaluating agent self-evolution on real business workflows with automated data pipeline to prevent contamination.
LLM capability for testing Terminal User Interfaces: benchmark across ratatui/Rust, bubbletea/Go, textual/Python showing 12% coverage in real applications.
CARE-Bench evaluates 11 LLMs on medical triage safety, testing when patient-facing models recommend escalation to clinicians.
AntiSkillBench: end-to-end benchmark evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for agents.
Pattern Completion Bias benchmark measuring how repeated UI patterns degrade multimodal LLM accuracy on screenshot-to-code fill-in-the-blank tasks.
LiveEvalBench: agentic, adaptive evaluation framework treating web generation as interactive problem with diverse valid implementations.
VetScore evaluates veterinary QA outputs by assessing citation faithfulness and weighting claims by medical harm potential.
Study compares GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 peer reviews on 300 ICLR submissions against human reviewer alignment.
Study evaluates zero-shot coordination robustness across independent algorithm implementations to assess practical agent alignment.
Systematic study reveals gender bias in LLM-based fake news detection across six state-of-the-art models using LIAR benchmark.
FOUND-AF benchmark evaluates ECG foundation models for atrial fibrillation detection under controlled leakage and deployment protocols.
DiagChain diagnostic benchmark evaluates LLM agents on staged attack chain reconstruction from security telemetry across 69 scenarios.
onepot-Bench 0 introduces proprietary benchmark for evaluating LLM abilities in wet-lab chemistry tasks, avoiding training-data contamination and measuring physical-world decision-making.
First systematic benchmark of Sheaf Neural Networks on inductive tasks, extending prior transductive-only evaluations across three diffusion mechanisms.
MedPRESS: 600-dialogue benchmark measuring patient-pressure-induced sycophancy in LLMs across medication, self-care, and triage scenarios.
SWE-Touch benchmark tests coding agents in shared workspaces with user code edits, exposing weaknesses in handling conflicting modifications during task execution.
ParEvalLayer detects biased partial evaluations of LLM agents, preventing premature benchmark conclusions from incomplete task runs.
Solution Hacking identifies shortcuts where LLMs achieve correct answers without valid reasoning on frontier science benchmarks.
Agentic Commerce World environment enables multi-agent evaluation with independent buyer/merchant objectives via Vibe Commerce Protocol.
Empirical analysis of why frontier LLMs fail at tabular prediction without fine-tuning, foundational insight for tabular models.
MonitrLLM: open-source evaluation infrastructure linking LLM conversation transcripts to user-defined task intent and outcomes.
Benchmark of 11 audio-language models and classifiers on closed-set sound source identification (2,242 clips, 23 classes).
CompressAgent benchmark evaluates reliability of compressed agent control contexts across Qwen models and task families.
Temporal replay framework evaluates enterprise agents against dynamic data across multiple moments within an episode, not just final state.
CallScreenBench evaluates on-device LMs as phone secretaries with adversarial callers and no oracle ground truth.
Human-written hallucination samples improve VLM benchmark stability across 4 languages vs. model-generated negatives.
MedUPS: benchmark of 21,874 mid-stream clinical decision points from 5,535 cases for LLM diagnostic assistance on uncommon medical scenarios.
Capability-taxonomy-driven pipeline for curating regression eval sets across multi-customer agent-extensibility platforms under query budgets.
LLMs as examiners silently omit valid answers when authoring test sets; greedy one-shot generation fails to enumerate complete solution spaces.
SCHEDBench: benchmark with 1,132 scheduling instances to evaluate LLM constraint faithfulness under natural-language variation.
Benchmarking PPG-based sleep staging datasets and metrics, showing substantial performance gap vs. EEG and advocating finer temporal resolution.
Calibration breaks under unseen subtype shift: models maintain accuracy but become systematically overconfident on novel fine-grained categories.
UpliftBench reveals metric disagreement, not model disagreement, drives uplift estimator ranking variance across 7 dataset families and 12 methods.
FinHardBench evaluates LLM-generated FPGA hardware for financial trading (latency-critical 5-10ns); 33 tasks span module generation, pipeline tuning, adaptation.
OpenAI's internal Astra model solved ten decade-old math problems for under $2K, matching Anthropic's cryptographic findings with Claude.
DeepSeek releases V4-Flash-0731, a 304B parameter model with enhanced agentic capabilities, outperforming larger competitors at $0.14/$0.27 per million tokens.
Simon Willison releases smevals, an open eval framework for benchmarking models, prompts, and inference harnesses across configurations.