Should you trust ChatGPT with your entire workflow?

Max Haining

Max Haining

12 Jul 20265 min read

Article cover image.

It’s been a fairly chaotic week, even by AI standards.

Anthropic spent the first half extending access to Fable 5 while everyone tried to squeeze in a few more builds before hitting the limits. I wrote about the 5 non-coding things worth doing with it and you now have till July 19th (allegedly) to play with it.

Article image.

Then OpenAI took over.


Within two days, it released GPT-5.6, ChatGPT Work, and GPT Live, its new voice model. Depending on who you read, this was either the moment agents finally became normal work software, the best voice product anyone has shipped, or an impressive collection of names and settings nobody fully understands yet.


Peter Yang tested the releases for a day and gave GPT-5.6 the most useful review I saw: “It’s got that dog in it”.


It keeps going. It stays with difficult work. And it seems less likely to hit a complicated task and decide you probably didn’t need it done anyway. But Peter’s praise came with a pretty long list of confusion too:


What exactly is Work versus Codex? Why are some things chats and others tasks? When should you use Sol, Terra or Luna? Why are there several effort settings before a normal person has even started the job?


Right now the pieces are impressive and still a little scattered. OpenAI may have made working with agents feel more mainstream this week but hasn’t fully made them feel simple yet. So let’s try to make sense of it today.


Window into the Future


ChatGPT Work is where this gets much bigger than the new voice model.

You can hand it a project, give it access to the files and apps you choose, and leave it working across the messy middle: research, analysis, documents, spreadsheets, presentations, websites. OpenAI says it can stay with complicated jobs for hours, splitting them into smaller steps and completing them without waiting for you after every click.

It’s basically the useful parts of Codex brought into the ChatGPT product almost everyone already understands well enough to open.

For non-technical people, that matters more than a benchmark score. You dont need to know how the agent is assembled. You can hand it the folder.

Article image.

There are loads of examples in OpenAI’s launch post: complete slide decks, financial models, documents that follow an existing company format, interfaces it can inspect and refine instead of generating once and abandoning.


That part looks excellent.


But the most useful criticism I saw came from Ethan Mollick a few weeks ago, then again after testing ChatGPT Work this week. His point is that these products still think like software tools.

In software, the codebase can serve as a source of truth. You can test whether something works. Changes are tracked. Failed branches, old versions and decisions are usually recoverable if someone cares enough to look.

Knowledge work is less cooperative. A polished report might contain 20 sources. It probably won’t tell you which two sources caused the author to change their mind, or that one important number came from a spreadsheet somebody described as “rough but basically right” in Slack.

A lot of the work is hiding around the artifact. Which means an agent can produce something that looks finished while missing the part that made the work make sense.

I keep thinking about my own newsletter archive.

There are dozens of published issues in there, plus screenshots, links, messy drafts, subject lines I dropped because they were repetitive or just a bit boring. If I ask ChatGPT Work to analyse it, I dont need a summary telling me we write about AI and the future of work.

There are dozens of published issues in there, plus screenshots, links, messy drafts, subject lines I dropped because they were repetitive or just a bit boring. If I ask ChatGPT Work to analyse it, I dont need a summary telling me we write about AI and the future of work.

That’s closer to actual knowledge work.

The models can now work for hours. They can inspect more files than I could read in a week, keep several routes open at once, build the deck, fix the deck, then build a tiny website for the deck because apparently we’re doing that too.

I still have the same brain I had last Sunday.

Article image.

The launch videos tend to end when the work is done but normal work doesn’t. And that’s a very different problem from “can the model do the task?” It can.

Can the rest of us keep up with what it did, understand enough of the route, and spend our attention in the places where a quick approval could turn into a very expensive afternoon? That’s the part I think teams will run into next.

There will simply be far more finished-looking work arriving than anyone is used to reviewing. We already have this with AI-written documents. Agents can multiply that asymmetry.

There’s a version of this that works brilliantly, obviously. The agent does the slow collection and construction, then returns the small number of decisions that actually need you. That’s what the good examples are moving toward.

But I want the receipt attached. So the first thing I’m trying with Work is a very unglamorous instruction at the end of the task:

Alongside the final output, give me a one-page work log showing:
- the sources you relied on most
- anything you ignored or could not access
- the assumptions you made
- decisions you made without asking me
- anything that still needs a human to check

I just want to know where to look before I approve something built from 70 files and three years of context.


This is also a lot of what we’re working through with teams at 100 School. Once the tools can do more, somebody still has to agree what needs checking, what needs recording and what “finished” means.

How to AI: GPT-5.6 + Work Edition 🤖

Here are a couple of ways you could use GPT-5.6 at work this week that are worth your time:

See you next week!

Share this

Ready to build real AI capability

Get weekly insights on adoption, team rituals, and practical workflows — straight to your inbox.

We respect your inbox. Unsubscribe anytime.

Person holding a lightbulb with an idea.