Think Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking
MERIT-Rank uses multi-trajectory reasoning with LLMs to improve document reranking robustness beyond single reasoning paths.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
MERIT-Rank uses multi-trajectory reasoning with LLMs to improve document reranking robustness beyond single reasoning paths.
AdaRepair-Mem adaptively balances episodic memory retrieval for LLM-based repository-level program repair across imbalanced datasets.
Unsupervised LLM safety detection via sparse activation anomaly detection without labeled unsafe data.
FedQoS enables asynchronous federated learning for multimodal sensors in smart vehicles with heterogeneous network constraints.
IBM Granite 5.0 TurboCTC: 470M-param Conformer ASR model with strided convolutions and Muon optimizer optimization.
Cooley law firm deployed ChatGPT Work to build GO Public, an IPO workflow tool for surfacing legal issues and automating document review.
SETTer applies sparse-encoder Transformers to long-horizon multivariate time-series forecasting with oversmoothing mitigation.
MATCH uses model-aware curriculum scheduling and hierarchical reward gating to improve LLM tool-use via reinforcement learning.
Conceptual agentic AI architecture for multi-domain decision support in Brazilian military command-and-control operations.
On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room" to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup's systems - all without OpenAI finding out about it for more than a week. No one in the war room was surprised; this w...
Astronex-World 1.0: open video world model with text/image conditioning, camera trajectories, and block-causal attention for real-time generation.
Framework translating EU AI Act technical requirements into 43 machine-checkable compliance criteria for generative AI systems, addressing seven gaps in original regulation.
Study of 67,200 responses from GPT, Claude, and Gemini across 112 languages shows geopolitical bias variance in LLM answers about Ukraine war.
Interview on consumer AI adoption and foldable phones; limited technical depth for frontier AI professionals.
AVTrace benchmark suite with 34,114 examples diagnoses temporal reasoning in multimodal models across audio-visual synchronization and event localization.
GS-Power-UCT improves Monte-Carlo tree search efficiency in stochastic MDPs by sharing states across trajectories with convergence rate O(n^-1/2).
MaSCoD multi-agent framework uses LLMs for causal discovery by organizing candidate variables before edge judgment; shows dataset-dependent performance on Auto-MPG, DWD, Sachs.
Stringological sequence prediction algorithm with Arithmetic Repetition Complexity measure achieves quasilinear time and polylog space for structured sequences.
Analysis of ReLU neural network approximation bounds in Sobolev norms for width W, depth L, and path-norm-constrained weights.
Graph hypernetwork framework amortizes physics-informed neural networks across PDEs by encoding operator relationships and enabling meta-training for reusable solvers.
Study of self-replicating neural cellular automata with mutation-driven diversity metrics to quantify emergent phenotypic and genotypic ecosystem organization.
TRACE framework enables training-free agentic retrieval over OCR-degraded historical archives for source discovery with accountability in political discourse analysis.
VākQA introduces a 2,001-pair Telugu spoken QA benchmark with speech audio and validates LLM-based evaluation methods against human judgments.
Yegge shuts down Gas Town project; Databricks raises Astra vector DB costs 60%, prompting reality checks on AI industry hype and unit economics.
Treble's voice simulation platform is used by voice AI model developers, AI wearable, and robotics companies
This session will explore how early-stage companies are building teams where humans and AI agents work alongside each other — and how founders can do that without sacrificing speed, accountability, or culture. Learn more at TechCrunch Disrupt 2026. Register before September 25 to save up to $200.
Since Specs' debut earlier this year, Snap has clearly been looking for an opportunity to explain why the smart glasses deserve to exist.
OpenAI launches Astra for Law, a domain-specific application with workflow customization, data integration, and compliance controls for legal firms.
Datasette 1.0a40 adds background task API, migrates to httpx2, fixes bugs ahead of stable release.
Datasette 0.65.5 patches security flaw allowing table permission bypass via trailing newline in names.