Vol. I · No. 120MON, AUG 17, 2026
Archive

The Archive

Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.

Learning When to Think While Listening in Large Audio-Language Models

Recent advances in Large Audio-Language Models (LALMs) have made real-time, streaming spoken interaction increasingly practical. In this setting, reasoning quality and responsiveness are tightly coupled: delaying reasoning until the speech endpoint can improve answer quality but moves deliberation into user-visible response delay, while answering too early risks committing before decisive evidence arrives. We introduce a learnable wait-think-answer control formulation for LALMs. Motivated by the incremental nature of human conversation, the controller decides under partial audio evidence when...

·

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we evaluate six cognitive tasks across three score levels: task, domain, and global levels. We compare hand-crafted acoustic features with self-supervised learning (SSL) embeddings. Results show that although SSL representations generally outperform hand-crafted features at lower levels, this trend reverses for MCI classification. Furthermore, task-specific constraints influence...

·

MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lost-in-conversation (LiC) gap. We trace part of this degradation to self-contamination: intermediate assistant replies enter later context and carry early deviations forward. Motivated by this mechanism, we propose MAIGO, an on-policy self-distillation method that reduces this contamination using history-cleaned references from the model's own policy. For middle turns, MAIGO removes prior assistant replies while preserving the user-visible shar...

·

Microsoft Copilot Cowork Exfiltrates Files

Microsoft Copilot Cowork Exfiltrates Files The biggest challenge in designing agentic systems continues to be preventing them from enabling attackers to exfiltrate data. In this case Microsoft Copilot Cowork (yes, that's a real product name ) was allowing agents to send emails to the user's own inbox without approval... but those messages were then displayed in a way that could leak data to an attacker via rendered images: Because these messages can contain external images that trigger network requests to external websites, data can be exfiltrated when a user opens a compromised message sent ...

·

Opus 4.7 Often Assumes a Military Audience

This is especially prevalent in Claude Desktop and less obvious in Claude Code. * Often leads long assets with a BLUF (bottom line up front), specifically military-adjacent terminology. * Has slipped into language where it refers to civilians several times, as though it means someone other than us * References to things like deep-dive, vision-intent framing, force multipliers, and reframes are more subtle because they do cross over but their prevalence is unmistakable. This isn't a complaint, it's more a note that, "hey guys, we see your training material in our outputs" and its interestin...

··

The end. What have I done

It seems to be working so far but I think I should have done this in GitHub

··

OpenMOSS-Team/MOSS-TTS-v1.5 · Hugging Face

# MOSS-TTS-v1.5 **MOSS-TTS-v1.5** is continued from [MOSS-TTS 1.0](https://huggingface.co/OpenMOSS-Team/MOSS-TTS). It preserves the main 1.0 capabilities, including zero-shot voice cloning, long-form speech generation, token-level duration control, Pinyin/IPA pronunciation control, multilingual synthesis, and code-switching. For the full 1.0 feature walkthrough, input schema, decoding hyperparameters, and evaluation tables, please refer to the [MOSS-TTS 1.0 README](https://huggingface.co/OpenMOSS-Team/MOSS-TTS). Compared with MOSS-TTS 1.0, v1.5 focuses on the following improvements: * **St...

··

Quoting Paul Graham

A lot of the emails I get from founders are now written in a hard-hitting journalistic style. I know they're written by AI, because no founder ever wrote this way before. And once you realize something is written by AI, it's hard not to ignore it. I have never knowingly finished reading an email signed by a human but written by AI. It feels like being lied to, and who would stand for that? [ ... ] It makes me think less of the author. It means they can't write well unaided (or feel they can't), and that they're trying to trick me. It's not impressive to use AI to write stuff for you; any teen...

·

Rethinking organizational design in the age of agentic AI

Amid rapidly growing adoption of enterprise-level AI agents, there’s a disconnect emerging between ambition and execution. Although 85% of organizations say they want to be agentic within the next three years, 76% say their current operations and infrastructure can’t support that change. They cite a lack of readiness across people, processes, and workflows. The sticky…

·

Sundar Pichai on AI, the future of search, and what’s happening to the web

Today, I’m talking with Google and Alphabet CEO Sundar Pichai, in a conversation we recorded just after the Google I/O developer conference. This is the fifth year Sundar and I have sat down after I/O, and it’s become one of my favorite Decoder traditions. There’s always a lot of news at I/O, and this year was no exception — Google has powerful new Gemini models, it’s putting AI agents in everything, and it’s making huge changes to Search on both the web and YouTube that will once again reshape the information ecosystem. That’s a lot to talk about, and Sundar and I got into all of it. But I a...

·

AI quietly turned HTML into a real alternative to PowerPoint and Word for client-facing docs. The blockers that made it impractical a year ago are falling one by one.

A year ago, generating a polished document as HTML instead of a PPT or a Word file was a fun idea with too many practical problems. Lately I've noticed every one of those blockers either gone or close to gone, and I've quietly stopped reaching for Office on a bunch of deliverables. Curious if others are seeing the same. **The blockers, and where they stand now:** **Design**. The old objection was "AI HTML looks generic and amateur." That's basically solved if you give the model a design skill or a style guideline once. You get consistent, on-brand output that looks more like a designed page...

··

Okay 27B made me a believer

I previously hated on this model, but I have just been impressed by it, and I understand the hype now. I have been working on a HTML5 game console and I decided to see if Qwen3.6 27B can handle making some quick games in it to showcase functionality (save games, console API handling for stat tracking and heartbeat management, meta data for the game, etc) I gave it 3 files, explaining how the API works, the gamepad controls, and a typescript shader for it to apply. Then I just game it a very simple prompt "make a breakout game for this console, in the working directory are reference files on...

··

EngineAI shared a view of its Shenzhen Intelligent Manufacturing Base, claiming an output of one humanoid robot every 15 minutes - that's 35,000 humanoid robots per year, the highest production rate publicly claimed by a Chinese humanoid robotics company

This is the highest output rate announced. besides this EngineAI has another aditional Zhengzhou 10K/year line planned. Its more than from what Leju Robotics, AgiBot, Unitree Robotics, and others have claimed for their humanoid robots per year. So its likely, China ready to output 100K humanoids robots per year.

··

Nobody wants to tell me why they only listen to their own Suno slop

Do you even like art? | Image: Cath Virginia / The Verge, Getty Images There's this alarming trend in the Suno subreddit. People aren't just prompting AI songs; they're sitting around listening almost exclusively to their own slop. And in some cases, they proudly proclaim that they don't listen to music on traditional streaming platforms anymore - it's just AI all day. "Does anyone just listen to their own music now and not even music on Spotify anymore.?" "I definitely listen to my own music most of the time now. Why wouldn't I? It's album after album of bangers" "Guilty as charged. It's an ...

·

AI warfare is already here

The Convention on Certain Conventional Weapons, an international forum that focuses on lethal autonomous systems, is hosted twice a year at the United Nations in Geneva. When Branka Marijan attended in November 2017, she thought the five-day sessions - which dealt largely in hypotheticals, speculating on a world where warfare was fought with killer robots - would be business as usual. After all, this was technology some thought might never be developed, and likely never deployed. That year, she quickly realized, was different. That distant, imagined future was suddenly closer and realer than ...

·

I let Claude rank every YC Spring 26 startup — round 2

Follow-up to my W26 post a few months back. Ran the same Claude pipeline on the YC Spring 26 (X26) batch. Same setup: for each company, Claude scrapes founder LinkedIn profiles, searches for press and traction signals, and checks the product to see if something real exists or it's just a landing page. Then it scores on founder credibility, product reality, market opportunity, and competition, and assigns a tier from S to D. Demo Day is June 16, so the batch is mid-flight and rankings will keep shifting as more companies launch. Most are B or C tier, which feels about right for this stage. ...

··

[D] Where do you go for serious AI research discussion online? [D]

Looking for communities where people actually dig into ML/AI research, not hype, not "look what I built with an LLM API," but discussions about papers, training dynamics, debugging real models, infra problems, that kind of thing. I'm specifically interested in places where you can post something like "I'm seeing X behaviour in my SSL training, here's the loss curve, anyone seen this before?" and get thoughtful replies instead of generic advice.

··
30 stories