Superhuman’s new auto-draft feature almost makes me like AI replies
Superhuman’s latest AI email drafting feature is its most convincing yet, generating replies that often required little to no editing in our testing.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Superhuman’s latest AI email drafting feature is its most convincing yet, generating replies that often required little to no editing in our testing.
Directional constraints improve exploration efficiency in safe reinforcement learning for robotics, balancing safety guarantees with task performance in constrained optimization.
Encoder-decoder transformer optimizes quantum circuits for fault-tolerant computing by minimizing T gates while maintaining functional equivalence.
Learning-accelerated ADMM algorithm accelerates scenario-based model predictive control for real-time planning via parallel computing and Moreau envelope learning.
Multi-task facial emotion recognition system for ABAW challenge using frozen lightweight extractors with temporal smoothing and ensemble fusion.
LLMs fine-tuned to predict chemical reaction mechanisms step-by-step, reducing hallucinations versus name-reaction prediction approaches.
Analysis of length bias in multiple-choice benchmarks shows length normalization over-corrects; proposes Bayesian alternative scoring for fairer ranking.
Constraint-aware federated RL aggregation method for microgrid energy coordination that prevents unsafe global behavior via penalty-based weighting.
Philosophical analysis of deploying opaque AI systems, examining user judgment, virtue, and autonomy-opacity tradeoffs in safety and control.
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision without explicit mechanisms for geometric consistency, leading to spatial hallucinations such as duplicated structures and misaligned geometry. These issues become more severe in 4D generation, where maintaining consistency across viewpoints and temporal evolution introduces additional challenges, including jitter, identity flicker, and structural drift. We present \textbf{Hallo4D}, a unified and model-agnostic framework for mitigating spatiotemporal halluci...
Farm site discovery from satellite imagery is a spatiotemporal candidate ranking problem because farm evidence is distributed across pasture, field boundaries, roads, buildings, and seasonal vegetation patterns. Direct farm labels are often incomplete, which makes fully supervised detection difficult. This paper proposes a weakly supervised pipeline for ranking dairy farm candidate clusters from seasonal Sentinel imagery and open map priors. The method uses aligned spring, summer, and autumn image tiles from County Cork, Ireland, with spectral bands, vegetation indices, built area indices, an...
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approaches either rely on prompting-based pipelines, which are computationally expensive, or require post-training on failure trajectories with step-level error annotations, which are costly to collect and difficult to scale. We argue that a practical failure attribution model should be lightweight and trainable without step-level supervision on failure data. To this end, we address unsupervised failur...
ESFP benchmark measures whether LLMs distinguish and coherently shift between neutral attribution and self-stance epistemic registers in contested claims.
Empirical study of 188 grokking runs shows representational priors must match task-relevant feature families to enable generalization; label-free invariance priors work via commutation symmetry.
Elenchos framework evaluates abductive reasoning in LLMs via formal-system mutation detection, exposing gap between pattern recognition and latent-hypothesis inference.
Probabilistic load forecasting framework for smart buildings handles input uncertainty via reconstruction and calibration of prediction intervals.
Anthropic pledges $10M to Canadian AI research initiatives, expanding regional R&D presence.
XGBoost emulation of high-dimensional likelihood functions in physics and cosmology improves efficiency over traditional global fits.
Bulkhead automates detection and remediation of container path-traversal vulnerabilities exacerbated by AI workload resource sharing (GPUs, agent workspaces).
FileMark VSCode extension uses line-anchored feedback to reduce token generation in Claude Opus (22%) and Sonnet (58%), cutting code-editing latency and cost.
Neuro-symbolic approach integrates MaxSAT constraint reasoning into vision-language models to enforce logical consistency in Sudoku solving.
Label-decoupled style augmentation improves domain generalization in multi-label remote-sensing classification by avoiding per-class contamination.
Cost-aware speculative decoding for Mixture-of-Experts LLMs optimizes expert activation patterns to reduce inference cost beyond token acceptance alone.
CARE-PPO combines PPO fine-tuning with loss prediction to jointly train LLMs for accurate numerical estimates and calibrated confidence signals in quantitative prediction tasks.
SeRIn proposes architectural separation of modality-specific refinement from cross-modal fusion in multimodal LMs for sentiment analysis.
Method for panoptic symbol spotting in CAD floor plans using text-aided multimodal analysis for industrial digitalization.
Internet of Agentic Things (IoAT) framework integrates autonomous AI agents with IoT, cyber-physical systems, and edge computing for closed-loop orchestration.
Demis Hassabis, during a panel session at the World Economic Forum in Davos, Switzerland. | Image: Bloomberg via Getty Images Demis Hassabis thinks the world needs an AI watchdog with the power to hit the brakes if frontier models become too dangerous. Writing in a blog post, the Google DeepMind CEO and cofounder said the US should lead the initiative, arguing that the country is the best place to set global standards "given its economic and technical standing." The organization, which could resemble existing regulators like the Financial Industry Regulatory Authority, would be made up of lea...
Jetson-PI enables real-time onboard VLA model deployment on low-power Jetson devices via foresight-aligned asynchronous inference for robot control.
EG-VAR uses Lean 4 formal verification with tool-calling to ground LLM reasoning in attested evidence and kernel-checked inference chains, eliminating hallucination.