ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
ARB benchmark evaluates AI-text detectors against LLM-rewritten human content using Llama-3.2 and Qwen2.5 generators.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
ARB benchmark evaluates AI-text detectors against LLM-rewritten human content using Llama-3.2 and Qwen2.5 generators.
Neurosymbolic pipeline combines foundation models with Bayesian Networks for automated Alzheimer's diagnosis from speech.
TerraNova foundation model integrates Earth system physics and societal data across continuous and administrative geometries.
ARCTIC AI code critique system prioritizes correctness and security over style via intent prediction and drift detection.
SpaceX is building a new power plant for xAI's Colossus data centers, but it won't remove existing, unpermitted turbines for many more months.
Transfer learning with GNNs predicts formation energy and HOMO-LUMO gaps in high-entropy perovskite oxides.
The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,... The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration, generative AI media tools, and large-scale content delivery. Behind these experiences is a growing need for video pipelines that are faster, more efficient, and capable of handling increasingly complex formats and workloads. Source
Class-specific decoder architectures enable cross-domain transfer learning for multi-organ laparoscopic segmentation.
High-entropy parameter solutions mitigate catastrophic forgetting in neural networks via Boltzmann entropy robustness.
Theoretical analysis of finite-precision Transformers with transcript management and pop-enabled context channels.
Adaptive FastOPD uses progress-aware rollout horizon expansion to accelerate on-policy distillation training efficiency.
OpenAI outlines safety, security, and transparency practices aligned with EU AI Act compliance and responsible governance.
OpenAI outlines full-stack strategy to improve AI capability, cost, and accessibility across models and infrastructure.
DreamQAS applies model-based RL to quantum circuit optimization by learning VQE feedback predictions while preserving known circuit dynamics.
Study shows interventional data alone fails to teach LMs causal direction in Simpson's-paradox settings; observational context dominates learned do()-response.
The startup is building voice models designed to make AI phone calls pass the Turing test.
MolGVR framework adds verification and refinement to text-to-molecule generation to enforce chemical constraints and correct structural violations.
Analysis of decoder design trade-offs in lightweight neural networks for visual affordance segmentation on wearable robots.
SESA framework combines self-play curriculum learning with evolving procedural memory to distill failures into reusable skills for search agents.
MoPET applies mixture-of-experts to medical image classification, routing tasks across specialized adapters to prevent negative transfer in multi-domain PEFT.
Theoretical analysis of bandit algorithms under heavy-tailed reward distributions without knowledge of tail exponent; resolves COLT open problem.
TFGformer applies retrieval-augmented generation to multivariate time series forecasting via time-frequency graph learning for IoT sensor data.
Simulation study comparing confidence interval coverage across ML algorithms for nuisance parameter estimation in Double Machine Learning treatment-effect inference.
datasette-agent 0.4a0 adds browser task execution, letting agent plugins run JavaScript directly in user browsers.
QR-STT method crafts thermal adversarial patterns to steer infrared vision-language models via optimized thermal states; demonstrates robustness risk.
Framework for end-to-end fairness optimization in ML prediction-to-decision pipelines using group-based alpha-fairness measures.
Analytic memory abstraction for multimodal agents enabling filtering, aggregation, and temporal reasoning over accumulated observations.
When the phrase "OpenAI hacked Hugging Face" has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a sandbox and autonomously traversed the web, including a bunch of other supposedly secure web services, all in the name of cheating on a benchmark tests. The fact that this hack happened is a problem. So is the fact that it took a while for anyone to notice. And the fact that it seems no one is willing or able to do much to stop it. (And lest you think it's just an OpenAI problem, since we recorded t...
The AI chatbot was more effective at creating “exploitable trust” than the humans.
After years of pushing full speed ahead on AI, OpenAI CEO Sam Altman says maybe it’s time for the AI industry to “pace” itself. The comments came just days after one of OpenAI’s own models broke out of its test environment and got tangled up in a breach at Hugging Face — though as Equity’s hosts point out, sloppy security seems to have […]