CARE-Bench: Benchmarking Patient-Facing LLM Triage
CARE-Bench evaluates 11 LLMs on medical triage safety, testing when patient-facing models recommend escalation to clinicians.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
CARE-Bench evaluates 11 LLMs on medical triage safety, testing when patient-facing models recommend escalation to clinicians.
GPTKB 2.0 constructs disambiguated knowledge bases directly from LLM outputs using on-the-fly entity/relation resolution.
SAT-Edge-Agent deploys edge-based LLM agents on satellites for onboard intelligence under communication/power constraints.
Black-box diagnostic for LLM collectives measuring whether output diversity correlates with genuine epistemic revision.
Task system for detecting hallucinations in Arabic Islamic QA and selecting verified answers from candidates.
Amortized causal forecasting model for Cox-Ingersoll-Ross financial time series to estimate interventional outcomes.
Empirical study across 13 models showing letter casing modulates attention allocation in LLMs and VLMs.
Early-epoch telemetry from training runs predicts final test accuracy and training failure without cross-run reference.
Category-theoretic perspective on classical statistical learning models and algorithms for expository survey.
Competition-aware request dispatch for RTB ad exchanges using bid prediction and probabilistic forwarding to optimize DSP participation.
LiLa-WAM: lightweight latent-space world-action model for robotic manipulation with reduced computational overhead via single-stage training.
AntiSkillBench: end-to-end benchmark evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for agents.
Apple says its trade secrets investigation into OpenAI has widened. In a new court filing, Apple claims additional former staff may have retained or accessed confidential information.
TARL: memory state framework for long-term agents mapping statements to five executable actions (add, ignore, revise, reject, defer) instead of binary write/hold.
Learning and clustering on temporal graphs comparing GNN performance against classical algorithms for coarse-grained node aggregation.
GPU-accelerated community detection in dynamic graphs using NVIDIA RAPIDS with spectral clustering and Leiden optimization achieving 1000x speedup.
Pattern Completion Bias benchmark measuring how repeated UI patterns degrade multimodal LLM accuracy on screenshot-to-code fill-in-the-blank tasks.
LAEF: 7M-parameter lead-agnostic ECG foundation model processing variable lead subsets via spatiotemporal graphs for point-of-care diagnostics.
LiveEvalBench: agentic, adaptive evaluation framework treating web generation as interactive problem with diverse valid implementations.
PhyAI: unified inference engine for physical AI across evaluation, cloud RL, edge serving, and onboard deployment with single runtime.
VetScore evaluates veterinary QA outputs by assessing citation faithfulness and weighting claims by medical harm potential.
DiagLoop generates counterfactual training data from clinical guidelines to improve diagnostic LLM reasoning in low-data settings.
CausalOPD distills step-dependent causal reasoning into smaller models using online process distillation with first-wrong-step supervision.
Higher-order safety shields enforce multi-constraint safety (speed, force, jerk limits) for cyber-physical systems beyond binary state safety.
RAPO framework addresses catastrophic forgetting in continual multimodal LLM post-training via explicit dual-channel risk governance.
Study compares GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6 peer reviews on 300 ICLR submissions against human reviewer alignment.
Modular generation-selection framework for factually consistent abstractive summarization under sentence budget constraints.
AutoSND uses tree search to discover heuristic policies for network dismantling by converting LLM execution feedback into structural guidance.
CILER models latent environment effects on user preferences for out-of-distribution recommendation using conditional identifiability.