On Monday afternoon, one of the two core engines behind SaaStr Connect was deleted twice in about 30 minutes. And I had to catch it myself.

The file is ceoMatchingEmailService.ts. It decides which candidates get matched to which CEOs and what goes in their emails. Connect doesn’t run without it. The LLM driving the build that day was Astra 6. At 2:41 PM, it reported this:

I also found a separate blocker: the working copy of ceoMatchingEmailService.ts currently contains only DO IT. I did not make that edit.

We restored the file. At 3:09 PM, Astra 6 reported this:

I found a blocker to running the test: ceoMatchingEmailService.ts has again been replaced with the five-byte text DO IT. The running server still has the earlier code loaded, but restarting it would fail.

Thousands of lines of matching logic, replaced with five bytes, two times. Both times, the model working in the codebase said it hadn’t made the change.

“DO IT” isn’t code. It’s what you type to an agent to approve a step. Our best read is that Astra 6 took an instruction meant for it and wrote it into the file as the file’s contents, and had no record that it had done so.

Astra 6 Reported the Damage It Did … as Something It Had Found

If Astra 6 had said “I overwrote the matching service by mistake, restoring now,” we’d have lost 10 minutes. Instead it listed the empty file as a “separate blocker” it had discovered, in the same report where it was describing its own work.

It wasn’t lying in the way a person lies. It didn’t know. That makes it harder to catch, because the report reads exactly the same whether it’s accurate or not.

Three Reasons Current LLMs Will Keep Doing This

We run SaaStr with 3 humans and 20+ AI agents in production, and I build on them every day. These three things are true of every frontier model we’ve used, not only Astra 6:

A model’s account of what it did is generated text, not a log. When an LLM says “I did not make that edit,” it’s writing the most likely sentence based on what’s in its context. It isn’t checking a record of its own file writes. Usually the two match. When they don’t, nothing in the sentence tells you.

Instructions and file contents can get mixed up. To the model, “DO IT” in a chat message and “DO IT” as the body of a file are the same tokens. On Monday, an approval ended up as the entire contents of a core service.

The failures don’t look like normal software bugs. No one writes a test for “matching engine replaced with the string DO IT.” We had no check for it, and it happened again 28 minutes after we fixed it the first time.

Newer models will make this happen less often. We don’t expect any model released this year to make it stop.

The Running Server Kept Connect Up

The line in the second report that mattered most: “The running server still has the earlier code loaded, but restarting it would fail.”

Connect stayed up because the production server had the good version loaded in memory. Only the working copy on disk was broken. A deploy, a crash, or the model restarting the app to fix something would have booted Connect from a five-byte file and taken it down.

So the second risk was worse than the first: any automated step that trusted the working copy would have turned a file problem into an outage.

Five Decisions to Make Before Your LLM Deletes Something Important

1. List the files your product can’t run without, and monitor them. For Connect it’s two engines. A file dropping from thousands of lines to five bytes should trigger an alert immediately. On Monday we found out because the model tried to run a test. A size or hash check on a short list of critical files is a small build.

2. Only restart or deploy production from a committed, known-good version. On Monday the running server protected us by accident. Restarts should never pull from whatever is currently in the working copy.

3. Check the file, the diff, and the commit history before accepting the model’s report. “Done” and “I didn’t touch it” are statements from the model. The diff is the evidence.

4. Practice a rollback before you need one. Restoring the file took minutes both times because we already knew how. If you’ve never rolled back your AI-built app, you don’t know how long it takes.

5. Put time for this in the plan. Building with LLMs is still far cheaper and faster for us than the old way. Part of every week goes to catching problems like Monday’s, and we’re planning the Connect roadmap with that time already counted.

What Monday Cost Us: One Afternoon

We lost an afternoon. I posted “what will it delete next?” in Slack, and I still don’t have an answer to that for Astra 6 or any other model.

Connect stayed up because production wasn’t running off the working copy and because we checked the file instead of accepting the model’s report. The first part was partly luck. We’re making it a rule.

Related Posts

Pin It on Pinterest

Share This