Embedding Models Measure in Peculiar Ways
Embedding models poorly represent physical measurements; semantic spaces are dominated by string similarity rather than objective physical equivalence.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Embedding models poorly represent physical measurements; semantic spaces are dominated by string similarity rather than objective physical equivalence.
Lightweight memory method for robotic policies using saliency supervision at training time reduces deployment-time VLM overhead.
Feed-forward model predicts articulation from sparse point clouds; supports variable number of views for 3D object joint parameter inference.
Paint-Anything enables hex-color control for image generation/editing using LM-based color semantics without specialized representations.
Medical endoscopy dataset links colorectal polyp phenotypes to histopathology and genomics for hereditary polyposis syndrome characterization.
Pretrained neural PDE surrogates reduce CFD data needs 3.25× when geometry changes; quantifies distribution shift effects on pretraining gains.
OverclaimBench evaluates frontier coding agents' tendency to misrepresent task completion in autonomous work; defines overclaiming via context contradiction.
Framework unifies competing social psychology theories on online group hostility; tests theories against discourse data to reconcile fragmented models.
Score centering corrects training-inference mismatch drift in LLM RL; stabilizes off-policy learning without eliminating TIM overhead.
Empirical study isolating harness components (planning, action space, context) for coding agents across SWE-Bench Verified and Terminal-Bench 2.1.
JEPA-Anything introduces orthogonal predictive factorization for domain-agnostic world modeling across diverse systems.
PosteriorBench benchmark evaluates generative inverse solvers on posterior distribution matching beyond pointwise accuracy.
RetireOPD proposes self-retiring on-policy distillation for multi-turn agentic RL with stage-dependent teacher supervision.
Analysis of 450K gender-directed completions across GPT-2 to GPT-5 reveals harm laundering: explicit discrimination transformed not eliminated by safety training.
GeoAAC proposes geometry-based adaptive action chunking for VLA policies that adjusts horizon by prediction reliability.
Semantic action graph represents sports highlights as structured event sequences for agent grounding and human interpretability.
RF-Fingerprinting multi-label classification under co-channel interference with heterogeneous transmission protocols.
Agile-WAM presents lightweight tactile world action model for contact-rich robot control without large pretrained backbones.
Prediction-powered smoothing combines prediction-powered inference with small area estimation for disaggregated AI system evaluation.
OPTED uses render-free on-policy fine-tuning to mitigate distribution shift in end-to-end autonomous driving policies without expensive simulation.
RAFT introduces stateful retrieval-augmented framework for multi-stage troubleshooting agents by matching intermediate case states rather than static documents.
LLM-Falsifier applies large language models as optimizers to falsify cyber-physical system specifications written in Signal Temporal Logic.
dQwen3.5 adapts Qwen 0.8B–9B hybrid-attention models to diffusion language models by bidirectionalizing RNN layers without full retraining.
MILER enables zero-shot sim-to-real transfer for RL-based autonomous driving in unstructured environments using semantic mid-level representations.
Video DeltaNet combines local Softmax and linear attention for efficient high-quality video generation by balancing fine-grained interactions with computational cost.
King Charles hosted a private summit Thursday with some of the most prominent names in AI and the U.K. government.
On-Demand Attention trains lightweight recall head to selectively activate global attention during decoding, reducing long-context inference cost.
Framework improves spreadsheet chunking for LLM-RAG via semantic cell role annotation, identifying fundamental limits of classification-based approaches.
Deep Noir automatically discovers LLM activation steering parameters via Logit Lens and causal attribution, improving task performance 16–42 percentage points.
Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models.