SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving
SWE-Serve benchmark evaluates agents on production inference engineering tasks spanning model support, runtime execution, and public APIs.
Every story tagged with this topic, ordered by date.
SWE-Serve benchmark evaluates agents on production inference engineering tasks spanning model support, runtime execution, and public APIs.
EquivSVA dataset enables testing whether LLM-generated SystemVerilog assertions capture true behavior vs. implementation-specific details.
Study shows compile rate is unreliable metric for LLM code vulnerability repair; proposes change-aware evaluation across 350M–6.7B parameter models.
Serving infrastructure (Ollama, Gemma, Phi) confounds tool-use evaluation results; model behavior inconsistency stems from serving layer gatekeeping.
Benchmark compares transfer learning, active learning, and pseudo-labeling for label-efficient cloud classification on Ground-based Cloud Dataset.
Decision-only LLM judge scales evaluation cost-effectively to 0.36% of full-judge fees, matches state-of-the-art on preference tasks, flags uncertain escalations.
Semiotic framework evaluates NLG fidelity and coverage by measuring contextual meaning and discourse references between texts.
Benchmark evaluation of 12 generative crystal structure prediction models; template retrieval outperforms diffusion and latent-variable approaches.
GitScholar dataset uses GitHub engagement signals to predict AI research impact, complementing citation and content-based methods.
TriWorldBench evaluates embodied world models via tri-view video consistency (head + bimanual wrist cameras) across 50 manipulation tasks.
CHiME-9 ECHI challenge evaluates multi-channel speech enhancement from noisy cafeteria conversations using Meta Aria and hearing aid microphones.
EADC benchmark evaluates LLM compliance with AI laws and regulations, detecting implicit covert risks beyond static explicit compliance checks.
CQ4OE benchmark systematically evaluates LLM-assisted ontology generation from competency questions with fine-grained provenance and structural adequacy metrics.
Text-to-SQL conformal abstention certificates misreport risk due to lenient single-database oracles; Spider-Realistic multi-instance swap reveals 2.73–10.23 point gaps.
GameHorizon Suite benchmarks multi-horizon gameplay capabilities across diverse models with unified data covering visual understanding, planning, and action control.
DolphinBench evaluates agent memory through task completion with cost/latency constraints, mapping Pareto frontiers across long-context retrieval scenarios.
BackTrend benchmark for weak-signal prediction: recovering underrecognized problems and emerging methods from mature scientific topics.
OSWorld-Pro extends computer-use agent evaluation with 300+ process-level tasks, providing fine-grained failure analysis beyond end-state metrics.
Exposure accounting metric measures LLM reasoning over graph-grounded corpora by controlling for context leakage via copy-ceiling baseline.
G-NAC: unsupervised clustering via graph neural cellular automata achieves 0.7951 ARI across 73 tasks, matching Genie with linear scaling.
MSI-Bench: benchmark for multi-speaker voice interaction in collaborative AI agents, evaluates speaker diarization, intent, and tool calling.
TicTacBench evaluates coding agents on RTL timing closure, a critical gap in existing hardware design benchmarks.
Systematic benchmark of long-tail loss functions on single-cell foundation models reveals architecture-agnostic failures on rare cell types.
Comparative benchmark of graph database engines on LLM agent workloads isolates query planning and ingest costs across vendors.
OMST optimizes strata construction for A/B test stratified sampling using multi-way decision trees to improve statistical power.
Comparative evaluation of EmbeddingGemma and fine-tuned variants for semantic job-candidate matching via hybrid retrieval with reciprocal rank fusion.
Diagnostic audit of Bayesian graph alignment convergence on 240+ exact and larger graph pairs reveals marginal disagreement gaps.
ChemCLIR-Bench: multilingual cross-lingual IR benchmark for chemical patents across five languages from Google Patents and EPO.
Temporal hierarchy forecasting reconciles hourly electricity prices and intraday spreads, improving accuracy up to 19.7% for battery arbitrage.
Critique of general LLM rankings: benchmark saturation, data contamination, commercial bias, and task-specific evaluation gaps.
Chronologic benchmark measures LM accuracy on historical English contexts 1831-1930 via pairwise comparison against multiple ground truths.
CraftBench-UE deterministic benchmark evaluates coding agents on 70 Unreal Engine tasks across C++, Blueprint, and editor scripting without LLM judges.
BrainWideBench: benchmark for evaluating transfer learning across animals and brain regions on multi-region neural recordings.
Multi-hop retrieval failures cluster predictably; dense-only ANN scoring weak vs LLM-judge pipelines for confidence calibration.
Benchmark comparing world models' continual learning on compositional tasks, measuring knowledge retention vs. speed of adaptation.
QuranicMMLU benchmark evaluates generative AI on Quranic Arabic across phonology, morphology, syntax, semantics, and pragmatics.
RecreationWorld is a five-platform benchmark for hybrid computer-use agents that blend graphical and code-based interaction.
Benchmark for robot failure diagnosis: sensor evidence auditing determines when robots should ask humans vs. act autonomously.
Benchmark comparing HyFyDy and MuJoCo motion-imitation RL pipelines on musculoskeletal fidelity vs kinematic accuracy.
Benchmark framework evaluating explanation quality alongside accuracy for vision-language models in forensic face recognition tasks.
EnterpriseVal benchmark measures GenAI reliability, safety, and ROI on enterprise workflows to predict deployment success vs. cancellation.
Statistical feature pool with LLM-generated signals exceeds neural baselines on TSB-AD-U time-series anomaly detection.
Audit of medical vision-language models (BioMedCLIP, CheXficient, MedSigLIP) on chest X-ray TB screening reveals ranking instability across datasets.
Benchmark evaluates LLM cross-lingual understanding of Chinese internet buzzwords with cultural semantics and safety implications.
Spoken Wikipedia Presentation Corpus extends ASR datasets with LLM-generated slides for multimodal speech recognition evaluation.
PRISM-BN introduces 5054-instance benchmark for text-to-parameterized Bayesian Network extraction in neurosymbolic AI.
HATS-en benchmark shows WER poorly correlates with human judgment; BERTScore variants better track semantic fidelity in ASR.
Chinese Competitive Debating Dataset: 182 professional debate matches with 360 judge annotations benchmarking LLM argument-tracking and adjudication.
Tooth segmentation on dental radiographs: input resolution dominates segmentation performance across FDI taxonomy with 1,422 annotated panoramic X-rays.
Medical endoscopy dataset links colorectal polyp phenotypes to histopathology and genomics for hereditary polyposis syndrome characterization.