Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
IdeaGene-Bench evaluates scientific lineage reasoning and idea generation by modeling papers as typed genome objects with inheritance tracking.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
IdeaGene-Bench evaluates scientific lineage reasoning and idea generation by modeling papers as typed genome objects with inheritance tracking.
Mathematical analysis showing score-matching forward-marginal error does not guarantee numerical stability in discretized diffusion samplers.
MulTTiPop dataset provides 3.5 hours of pop music with aligned multitrack MIDI for automatic music transcription evaluation.
SLORR introduces a stateless in-training low-rank regularization framework for neural network compression without architecture modification.
Large-scale log analysis of 77,543 students using AI learning assistant Syntea reveals actual usage patterns in distance education.
Graph algorithms applied to UMAP's internal k-nearest-neighbor structure improve data manifold interpretation beyond 2D projections.
AUTOPILOT-VQA benchmark evaluates vision-language models on safety-critical incident reasoning from dashcam video.
ARDY enables real-time controllable 3D human motion generation via autoregressive diffusion with text and kinematic constraints for interactive applications.
Proposes persistent knowledge object model for LLM workflows combining symbolic forms, object identity, and live-image thinking for tool use and execution.
Shows quantization from 8-bit to 2-bit induces behavioral divergence in LLMs beyond accuracy/perplexity metrics; introduces correctness agreement metric.
Challenges Super Weights hypothesis by showing isolated training of critical parameters fails on OLMo models; selective parameter training ineffective.
Evaluates AMALIA (9B Portuguese model) on moral foundation coding; distinguishes agreement from validity for theoretically-grounded construct annotation.
BioModule: lightweight temporal transformer predicting biomechanical attributes from 3D pose estimates for rehab and sports science.
Models adaptive reasoning in control policies via autoregressive latent space organized as memory palace for iterative information retrieval.
Deep learning framework for joint narrowband interference cancellation and soft demodulation in OFDM systems with non-Gaussian residuals.
Proactive memory agent running alongside action agent mitigates behavioral state decay in long-horizon tasks via structured memory bank updates.
Multi-modal 3D terrain reconstruction for wildfire-prone regions using image-based methods with DEM priors as geometric guidance.
In its bid to spend less on GPUs from providers like Nvidia, Meta is on track to start making its the latest versions of its AI-specific chips in September.
Molecular dynamics (MD) simulations are among the most demanding workloads in computational science. Using them, researchers can observe atomic behavior in... Source
Graph RL agent optimizes liquidity placement on Bitcoin Lightning Network via message-passing policy and PPO with action masking.
Having proven how valuable compute can be, the company finds itself at the center of a market everyone wants to be in — while simpler technologies and less interesting companies get rich on the sidelines.
Study calibrates LLM judges for citation verification in deep-research systems, testing whether frontier models are necessary for rubric-based reward signals.
Windows 11 updates could soon include fixes for more security issues at once. Microsoft said in a blog post on Thursday that it's now using AI to "identify potential issues earlier," which means "customers will see a higher volume of security updates included in each security release." Hackers, even amateurs, have increasingly been using AI to quickly exploit security weaknesses over the past several months. Security researchers are also using AI to find issues faster, leading to more frequent high-severity vulnerabilities, like the "Copy Fail" exploit that impacted nearly every Linux distrib...
About two weeks after OpenAI's GPT-5.6 was caught up in regulatory drama - rolled out only to government-approved organizations during a "limited preview" period - the company has received the Trump administration's green light for a public rollout of the model. OpenAI CEO Sam Altman called it "the best model we have ever produced." To celebrate, OpenAI also unveiled a new AI agent on the same day: ChatGPT Work. It's billed as a combination of ChatGPT and Codex, allowing the everyday non-technical user to take advantage of Codex's capabilities for non-coding tasks, and it's powered by the GPT...
ProjAgent uses procedural similarity retrieval to improve repo-level code generation by matching functional logic across codebases.
Benchmark of training-free relaxed speculative decoding techniques shows speed-capability trade-offs for LLM sampling without retraining.
SolarChain-Eval benchmarks autonomous agents in decentralized energy markets with physics constraints and trustworthiness metrics.
Framework compares test-time resampling vs. rerouting for cost-aware LLM selection under realistic budget constraints.
WebSwarm orchestrates recursive multi-agent LLM web search with adaptive collaboration to improve search depth and evidence coverage.
EdgeRefine applies Jaccard sampling to balance privacy and utility in GNNs under edge differential privacy constraints.