llm-gemini 0.31
llm-gemini 0.31 tool adds support for Gemini 3.1 Flash-Lite, now out of preview; functionality unchanged since March.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
llm-gemini 0.31 tool adds support for Gemini 3.1 Flash-Lite, now out of preview; functionality unchanged since March.
The week leading up to Thanksgiving 2023 was the AI industry's biggest soap opera moment. OpenAI CEO Sam Altman was abruptly ousted from his role at the ChatGPT-maker. The explanation? That Altman was "not consistently candid in his communications with the board." Now, via witness testimony and trial exhibits in Musk v. Altman, the public is getting a concrete look behind the scenes of that dramatic weekend for the first time, much of it centered on former CTO Mira Murati. It was a unique situation in that the rollercoaster of a power play - which seemed to change every hour - took place, in ...
AirPods Pro 3 | Photo by Amelia Holowaty Krales / The Verge Apple's rumored AirPods with cameras are nearing a stage where the company will test early mass production, Bloomberg's Mark Gurman reports. Currently, Apple testers are "actively using" prototypes that are in the design validation test stage, which is one step before the production validation test stage. The AirPods' cameras "aren't designed" to snap photos or video but instead can take in "visual information in low resolution" that users can query Siri about, like asking the AI assistant what they should cook with the ingredients t...
**Postmortem: Alien Pinball — built with Claude + ChatGPT + Suno + LittleJS** Just shipped a browser pinball game. Short writeup of the AI workflow in case it's useful here. **The game** — Full physics pinball: multiball, an A-L-I-E-N rollover multiplier (caps at 5x), skill shots, escalating combos, outlane gutter saves, and a wizard-mode centipede boss you fight while juggling 3 balls. Browser, mobile-friendly, no install. Play it: [https://focaccai.itch.io/alien-pinball](https://focaccai.itch.io/alien-pinball) **Setup.** Claude Code Max, Opus model for the heavy lifting. Roughly half my...
Can Sam Altman—or any CEO—be trusted with super intelligence?
The developer of Firefox says it has "completely bought in" on AI-assisted bug discovery.
"We are going to be saying goodbye to the swipe," CEO Whitney Wolfe Herd said.
Reddit discussion on Claude Pro usage limit increases and whether they adequately address user constraints.
Simon Willison releases Big Words, a simple URL-based text-to-slide tool for his vibe-coded macOS presentation framework.
This is incredible research. I'm only halfway through the post but I'm already racing. Could I/an average person build a tool to help with a normal person using the findings? Could it be paired with one of Anthropic's earlier tools to identify the "emotions" Claude is feeling when it uses certain language, almost like a lie detector? Could we look at the patterns in the language when hiding misalignment and see if Claude falls back to certain syntax? Also, it's such an interesting addition to the 10 ft wall, 11 ft ladder problem. We can read its thoughts, but sometimes it hides its th...
Security warning: Open-OSS/privacy-filter model on HuggingFace contains obfuscated malware payload in loader.py that executes arbitrary commands.
OpenAI is launching an optional safety feature for ChatGPT that allows adult users to assign an emergency contact for mental health and safety concerns. Friends, family members, or caregivers designated as a "Trusted Contact" will be notified if OpenAI detects that a person may have discussed topics like self-harm or suicide with the chatbot. "Trusted Contact is designed around a simple, expert-validated premise: when someone may be in crisis, connecting with someone they know and trust can make a meaningful difference," OpenAI said in its announcement. "It offers another layer of support alo...
ActCam enables zero-shot joint control of character motion and camera trajectories in video generation via diffusion models conditioned on depth and pose.
UniPool proposes shared expert pool across MoE layers instead of per-layer isolation, reducing parameters while maintaining performance in production models.
BAMI mitigates precision and ambiguity bias in GUI grounding agents without retraining using masked prediction distribution attribution.
EMO trains MoE models for emergent modularity, enabling selective expert activation by domain without performance degradation in memory-constrained settings.
VHG generates valid, challenging math problems via three-party self-play with verifier feedback to avoid reward hacking in LLM problem generation.
Analysis of ~89K Arena comparisons across 116 languages reveals global LLM leaderboards are statistically unreliable due to strong heterogeneity in preferences.
Optimizer-model consistency: using same optimizer for LLM finetuning as pretraining reduces forgetting while matching or exceeding LoRA performance.
Framework for comparative LLM safety scoring without labeled benchmarks via scenario-based audits with instrumental validity chains across languages and sectors.
AI co-mathematician workbench enables mathematicians to interactively collaborate with agentic AI on ideation, proofs, and theory building with stateful artifact tracking.
Mozilla used Claude Mythos preview access to identify and fix hundreds of Firefox vulnerabilities, shifting LLM-generated security reports from noise to actionable findings.
Positive-only policy optimization with implicit negative gradients improves RLVR for LLM reasoning by avoiding coarse failure discrimination in GRPO-style training.
SIRA framework improves retrieval-augmented agents by modeling expert search priors, reducing retrieval rounds and latency for organizational knowledge bases.
Extension of Venn-Abers predictors to unbounded regression using conformal prediction; narrow technical contribution to probabilistic forecasting.
Chromophore-region 3D mechanism graphs for fluorescent protein quantum-yield prediction; domain-specific protein modeling task, limited AI audience relevance.
Comprehensive benchmark for Multimodal Domain Generalization revealing inconsistent evaluation protocols and fragmented research; validates real-world robustness challenges.
StraTA framework adds explicit strategy sampling to agentic RL, improving credit assignment and exploration over long-horizon LLM decision-making.
Hybrid concept-based and abductive explanations for vision models using formal causal reasoning; advances interpretability beyond single-concept limits.
GlazyBench dataset (23K formulations) for ceramic glaze property prediction and generation; domain-specific multimodal benchmark with limited broad AI relevance.