How Cursor deploys AI inside the enterprise
Cursor's Forward Deployed Engineers help enterprises implement AI agents as software factories, per Pauline Brunet.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Cursor's Forward Deployed Engineers help enterprises implement AI agents as software factories, per Pauline Brunet.
SpaceX reportedly showed investors a "handset-like" AI device before going public. It could be another signal SpaceX wants to expand into wireless.
The actor and investor is joining forces with Morgan Beller, who was previously a GP at NFX, to invest in early-stage startups.
Framework measuring gap between LLM-generated and human research ideas via reverse-engineered paper lineages and two-axis evaluation.
Layer-wise analysis shows single transformer layer recovers most RL post-training gains, challenging uniform parameter update assumptions.
Language-critique framework uses natural language supervision instead of scalar signals for imitation learning from suboptimal demonstrations.
AutoMem treats memory management as trainable skill, enabling LLMs to autonomously decide file-system operations and knowledge organization.
Theoria verification architecture converts solutions into auditable state-transition sequences with explicit justifications between formal proofs and opaque scoring.
State-prediction separation hypothesis splits transformer computation into distinct streams for token prediction and state storage, improving efficiency.
FurnitureVLA enables long-horizon bimanual furniture assembly via vision-language-action models with VR teleoperation and progress-enhanced training.
Audit exposes reliability issues in coding-agent benchmarks (GSO, SWE-Perf, SWE-fficiency): runtime instability, scoring artifacts, selection bias.
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, or risk being blocked by default on many publisher sites.
Cartridge distillation detects stealth biases in LLMs by exposing preferential signals in logit distributions invisible to text-based analysis.
TiRex-2 extends xLSTM-based time series foundation model to multivariate forecasting with streaming and covariate dependencies.
GPU-parallel linearization error bounds for robust optimal control of nonlinear and neural network dynamics enable real-time planning with certified constraints.
World from Motion reconstructs freely renderable dynamic 3D Gaussian scenes from monocular video via conditioned video models and synthetic artifacts.
Empirical comparison of quantum vs. classical ML across seven model pairs finds insufficient evidence for quantum ML computational advantages.
Constraint programming and scheduling method optimizes resource utilization in autonomous labs for metal-organic framework synthesis experiments.
Neural Certificate Pricing exploits polynomial verification in combinatorial optimization by training neural networks to predict dual prices for certificates.
Adversarial generator-discriminator framework augments RLVR with learned style discriminator to prevent diversity collapse and reward hacking in LM training.
QuasiMoTTo uses quasi-Monte Carlo sampling to reduce redundancy in parallel test-time inference scaling for language models.
Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer... Reinforcement learning (RL) is central to aligning language models, from reinforcement learning with human feedback (RLHF) within AI assistants to newer reinforcement learning with verifiable rewards (RLVR) workflows for reasoning and agent tasks. RL is now becoming a practical technique for specialized AI where enterprises need more accurate agents for domain-specific workflows. Source
Decision-aware training augments energy score objectives with differentiable decision loss to optimize sample-based generative models for downstream task costs.
Diffusion-GR2 uses block-diffusion decoder for parallel position decoding to accelerate generative reasoning re-rankers with minimal accuracy loss.
Learned 3D Gaussian representation compresses structured and unstructured volume data more efficiently than INRs via explicit geometry encoding.
Bilingual evaluation protocol for cross-lingual speaker verification on Iberian languages isolates language effects from speaker variation.
US lifts curbs on Anthropic’s advanced Fable and Mythos models.
Adversarial pragmatics benchmark for LLM safety evaluation addresses instruction conflict, embedded commands, and policy ambiguity beyond binary pass/fail.
AGC-Bench unifies 78 datasets spanning multiple domains into HELM-standardized artificial general creativity benchmark for LLM evaluation.