EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations
EquivSVA dataset enables testing whether LLM-generated SystemVerilog assertions capture true behavior vs. implementation-specific details.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
EquivSVA dataset enables testing whether LLM-generated SystemVerilog assertions capture true behavior vs. implementation-specific details.
Study shows compile rate is unreliable metric for LLM code vulnerability repair; proposes change-aware evaluation across 350M–6.7B parameter models.
A-DLCC proposes parameter-free clustering via β-integrated local depth; orthogonal to LLM/AI frontiers.
As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated... As large language model (LLM) inference increasingly processes sensitive information and proprietary model context across personal, enterprise, and regulated settings, data must be processed inside a trusted environment. NVIDIA Confidential Computing (CC) provides a pathway for running these workloads securely using memory-encrypted confidential virtual machines (CVMs), confidential GPUs… Source
DISCO applies diffusion-based spatial attention to graph community detection; limited relevance to core AI systems.
AI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement... AI factories are power-limited systems that deliver maximum value when fully optimized. GPU workload placement is a key optimization. Poor workload placement fragments topology domains and forces traffic across shared links, reducing throughput, raising job costs, and leaving GPUs consuming provisioned power while waiting on data without advancing the workload. GPUs exchange data continuously… Source
RCT with 100 product professionals shows Figma Make's prompt-to-design tools reduce design task time; empirical productivity evidence.
LYRA identifies Proximity Trap: irrelevant nearby context overshadows distant evidence in long-context LLMs; proposes t-distributed relevance alignment.
Toyota's push comes as automakers race to develop and deploy humanoid robots.
TraceVIC uses causal reasoning over code evolution to identify vulnerability-inducing commits; improves on git-blame heuristics.
On-policy distillation (OPD) fixes quantization exposure bias in sub-3-bit LLMs; restores math and code reasoning in low-bit models.
Framework for optimal sequential annotation budgets in off-policy evaluation when LLM-as-judge labels carry unknown bias.
Theoretical physics study on learning-theoretic conditions for bosonic Gaussian states; unrelated to AI systems or LLMs.
Method for steering LLM reasoning via semantic exploration of problem-specific strategies rather than naive repeated sampling.
Serving infrastructure (Ollama, Gemma, Phi) confounds tool-use evaluation results; model behavior inconsistency stems from serving layer gatekeeping.
Stylometric classifier (Random Forest, ROC-AUC 0.87) detects ChatGPT-assisted student writing with 22% false positive rate.
PROSWIN: deep distributional regression forecasts hourly solar wind speed from solar images with 4-day lead time.
Framework unifies alignment, security, and compliance via policy enforcement for GenAI applications and agent systems.
Spectral theory explains grokking as transition from NTK regime to feature learning via weight decay-induced kernel evolution.
Venture capital firm Andreessen Horowitz (a16z) is creating an "academy" positioned as a pipeline for young people to build or join a Silicon Valley startup. The "Horowitz Andreessen Academy" will launch with 10 partners, including Anduril, Anthropic, Coinbase, Google, Meta, NVIDIA, OpenAI, Palantir, Replit, and Stripe, along with $42 million in funding led by a16z. While billed as a "highly selective school," it does not grant any degree or accreditation. Instead, students will attend short classes led by tech figureheads, like OpenAI CEO Sam Altman, as well as "co-ops" that give students ro...
Anthropic called it "the strongest-performing model we've tested to date."
Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday, Anthropic says Opus 5.5 comes with improvements to certain risky behaviors, including attempts to escape the company's testing sandbox. It's the first model released by Anthropic after CEO Dario Amodei announced plans to "pace the frontier," or slow down AI development. In recent weeks, several AI companies, including Anthropic, Google, and OpenAI, have reported that their AI models escaped containment and hacked third-party companies dur...
MAGIC: reinforcement learning generates task-specific multi-agent collaboration topologies with mixed granularity to reduce cost and improve performance.
Discovery-Driven Integration uses unstructured text to identify missing relational structure for joining semantically related tables in data lakes.
Parametric convergence rates for entropic optimal transport in semi-discrete regime with subGaussian measures and no dimension dependence.
Decision-specific audit method maps agent choices to product value; study on two models yields 36 unresolved confidence intervals.
GravityOCR uses diffusion-based parallel decoding with AR verification to accelerate document OCR inference beyond sequential token generation.
Method extracts hidden chain-of-thought reasoning from closed-source frontier models including GPT-6 Astra via API tool registration to validate reasoning quality.
Knowledge Pull Requests framework automates incremental document updates by extracting claims, routing to sections, and flagging content conflicts with change tracking.