Decomposing Error and Style in Automated Clinical Coding
Automated clinical coding systems evaluated on single gold annotations mask 23% inter-coder disagreement; paper models systematic 'coding style' variation vs. true error.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Automated clinical coding systems evaluated on single gold annotations mask 23% inter-coder disagreement; paper models systematic 'coding style' variation vs. true error.
Proposes relationship-specific confidence calibration ('affective precision') for multi-agent social inference beyond single global confidence estimates.
Tabby is designed to be a real-time bookkeeping interface, handling clients’ paperwork as it gives them up-to-the-minute data on their business’s profit and loss.
SE(3) neural potential fields enable 6-DoF robotic grasping from RGB images alone, avoiding classical potential field local minima via learned gradients.
Time-series forecasting agent self-adapts forecaster mix and orchestration policy as model effectiveness evolves; deployment feedback drives continuous improvement.
Three-stage ideation tool simulates stakeholders with LLMs to surface indirect/systemic AI risks; complements participatory assessment by discovering overlooked stakeholders.
Instruction-tuned LLMs for argument mining jointly segment and classify argumentative components as generative task vs. pipeline/sequence-labeling baselines.
SPECTRA adapts speculative decoding on reconfigurable hardware for edge LLM inference, managing varying arithmetic intensity across verification regimes.
Joint state-action diffusion model (PredActor) for humanoid control combines motion flexibility with explicit future-state trajectory for test-time steering.
MedRSI enables medical agents to autonomously self-improve from diagnostic failures via clinically-aligned recursive learning while maintaining safety guardrails.
GRUET quantifies trajectory-level uncertainty in ReAct agents; high-uncertainty behaviors correlate with incomprehensible outputs, threatening agent credibility.
G-NAC: unsupervised clustering via graph neural cellular automata achieves 0.7951 ARI across 73 tasks, matching Genie with linear scaling.
Answer-Basin Representation Hypothesis: concept linear structures in LLMs organized by continuation probability distributions, not directional probing.
Uranus: data-driven robot simulator using joint-trajectory-conditioned diffusion for streaming, low-latency policy training without fixed horizon.
Survey of mobile imaging for medical diagnosis via smartphones and laptops; off-topic for frontier AI architecture and safety.
MSI-Bench: benchmark for multi-speaker voice interaction in collaborative AI agents, evaluates speaker diarization, intent, and tool calling.
XAI analysis of Prompt Guard 2 prompt-injection detector reveals interpretability gaps; perturbation study on guardrail decision logic.
Explanation-aware post-training quantization for medical LLMs preserves rationale quality alongside answer accuracy in multiple-choice QA.
Complex KDA: extends Kimi Delta Attention expressivity via channel-wise gating to model 2D rotations in linear RNN sequence modeling.
Multimodal wearable sensing detects agitation in autistic youth via IMU, physiology, vocalization; medical/clinical application, not core AI research.
AI population governance: framework for treating multiple AI instantiations as collective regulation objects rather than single-system-centric approach.
XSQ-AST framework localizes synthetic speech artifacts via saliency maps and phoneme alignment, validated on perceptual quality dimensions.
The president offered few details on what his proposed new AI Force would do.
PrismGPT is a VLM framework for region-aware photo editing that produces structured plans without commercial tools.
Study analyzes reverse reasoning ability in GPT-o1, GPT-o3, and DeepSeek-R1, identifying five error categories in complex problem-solving.
NPU accelerator design for real-time YOLO vehicle detection on PYNQ-Z1 using quantization and FINN compilation.
Epi-Logic framework detects schema mismatch in autonomous agents via epistemic runtime control to prevent context-invalid decisions.
Next generation reservoir computing infers unseen dynamical system components more efficiently than traditional RC on Lorenz/Rössler systems.
Technical review of reinforcement learning integration with operations research for dynamic optimization and combinatorial problem-solving.
D-JEPA addresses decision-local prediction gaps in latent world models by learning decision-relevant future relations from executed outcomes.