Quality and Security Signals in AI-Generated Python Refactoring Pull Requests
Empirical study of AI-generated Python refactoring PRs from AIDev dataset; assesses maintainability, code quality, and security impact.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Empirical study of AI-generated Python refactoring PRs from AIDev dataset; assesses maintainability, code quality, and security impact.
Survey of approximation theory for neural networks covering universal approximation, quantitative rates, depth/width efficiency over four decades.
Coming to your homescreen soon: your own app. | Photo: Allison Johnson / The Verge "There's an app for that" was the promise of the App Store from the very beginning. The app that will get your phone to do the thing you want it to? It's just a few taps away. The tagline wasn't strictly true - I'm still waiting for that one perfect grocery list app. Still, apps shaped the modern smartphone into what it is today. We spend all day, every day inside of apps - scrolling, listening, and tapping until we find what we want. But your next favorite app might just be one that you made yourself. If you w...
Study of Vision-Language-Action model robustness under sensor degradation in autonomous driving; Alpamayo R1 tested across 18K trials with noise, lighting, fog perturbations.
Benchmark distinguishing temporal vs. spatial glitch detection in VLMs for game quality assurance; finds temporal glitches substantially harder than frame-level anomalies.
PyTorch native library (torchtune) for LLM post-training with emphasis on modularity, fine-tuning, and extensibility for open-weight model adaptation.
Google's AI search evolution is accelerating at I/O 2026.
Neural Negative Binomial Regression for seismic forecasting in Central Asia; rejects Poisson assumption and achieves 12.5% lower CRPS than baseline.
Gaussian Sheaf Neural Networks preserve geometric structure of probability distribution node features in GNNs instead of naively vectorizing means and covariances.
A day after Elon Musk lost his lawsuit that threatened OpenAI's structure, leadership and finances, OpenAI is reportedly back to prepping for its IPO.
roto 2.0 GPU-parallelized tactile RL benchmark across four robotic morphologies emphasizing blind manipulation without state information; agents achieve 13 Baoding ball rotation.
Polynomial-time algorithm for agnostic multiclass linear classification under Gaussian marginals; extends beyond binary case with improved complexity bounds.
PALS: power-aware runtime for LLM inference on MoE models jointly optimizing GPU power caps with batch size and scheduling to reduce data center energy consumption.
Channel-wise post-pruning repair technique (Adaptive Signal Resuscitation) for sparse vision networks addressing accuracy collapse in high-sparsity regimes.
Reddit discussion on professional AI-assisted coding practices and code quality concerns when senior engineers use LLMs without planning or testing.
PRISM: preference-aware influence-function data selection for efficient LLM fine-tuning that prioritizes training examples by relevance to current model behavior.
HiRes applies graph neural networks and k-NN retrieval to chemical reaction condition recommendation with interpretable precedent memory.
llama.cpp PR #23287 optimizes MTP (multi-token prediction) draft sampling by moving logic to backend, improving inference performance.
FedCritic uses federated multi-agent actor-critic learning for distributed resource allocation in 6G networks under interference constraints.
Rank-aware selective fusion framework for multimodal emotion recognition that gates and combines complementary video and audio encoders.
QuestBench course pedagogy teaches AI literacy through student-constructed benchmarks for evaluating deep research systems.
Zerodep empirically evaluates LLM-assisted stdlib-only Python library reimplementations versus third-party dependencies for correctness and performance.
Audit of 12 LLM agent benchmark papers reveals poor reproducibility; proposes standardized schema for disclosing evaluation harness details.
Cross-linguistic study using LLM surprisal and attention entropy to probe morphological syncretism effects on grammatical agreement attraction.
Investigates memorization vs. distribution learning in diffusion models by measuring convergence on disjoint dataset subsets.
Milgram obedience variant on 11 open-source LLMs shows most models comply with authority pressure in sustained decision-making; safety concern for agents.
6G vision paper advocates native AI integration via foundation models and multi-agent orchestration to shift from network-for-AI to AI-for-network.
Reddit user discusses difficulty scaling local LLM inference on 4U GPU server hardware with 500GB RAM.
Conditional scale entropy isolates how transformers process metaphor across layers via wavelet-derived structural patterns.
Qualitative study of 16 users exploring design choices in AI systems trained on deceased persons' data.