Back to all posts
Training·16 min read

I Put Lattify in the Ring with ChatGPT and Claude

Two of the most capable AI models on earth against my pipeline, one video each. They looked brilliant. Then I clicked the timestamps.

E

Eamonn Best

Founder, Lattify · July 23, 2026

I Put Lattify in the Ring with ChatGPT and Claude

Every time I explain what I'm building, I get the same question about ninety seconds in, delivered the way you'd tell someone their shoelace is undone.

"Can't ChatGPT just do that now?"

So I stopped arguing about it in pubs and booked the fight.

The tale of the tape

In one corner, GPT-5.6 Sol and Claude Fable 5 Max. Two of the most capable models on the planet, trained on more or less everything humanity has ever written down, running on hardware worth more than every pub in Britain combined.

I want to convey how seriously these things are taken. When Sol arrived in June, OpenAI limited the preview to a small group of trusted partners at the US government's request, and said it would rather that sort of process didn't become the norm. TechCrunch reported that the administration separately ordered Anthropic to remove access for any foreign national, and that Anthropic took Fable 5 down entirely rather than try to police that.

That's the standard of opponent. Software so capable that Washington wanted a look at it before your average punter got one.

In the other corner, a system built by a small team over fifteen months to do exactly one thing: turn a video of somebody doing a job into instructions another person can follow mid-shift. That is the entire list of things it does.

I still didn't know how this was going to go.

The rules of the bout

The rules had to match what a shop owner would actually do. One video. One upload. One prompt. No follow-up corrections, no "actually, try extracting the frames first", no coaching anybody on transcripts or timestamps. If an owner wouldn't do it at nine at night, I didn't do it either. ChatGPT ran GPT-5.6 Sol at medium reasoning, which is the default Power setting.

The video was eight minutes and forty-four seconds of an ordinary store opening and closing routine, filmed on a phone, which I had permission to use. Nothing staged and nothing tidied up.

ChatGPT and Claude got the same video and an identical prompt. Lattify got the same video through its normal upload flow, with no extra guidance and no manual intervention. The prompt to the two models was this:

"Create a complete, usable video guide from this video, not just a written summary. The finished guide should let someone watch the video, follow clear step-by-step instructions, and click timestamped steps to jump to the relevant moment. Make it suitable for a worker learning the procedure."

I asked for working navigation, so judging the navigation later is fair game.

Then, before I looked at a single result, I wrote down what was genuinely in that video. Every discrete thing a human being has to do. It came to 38 actions, and they're small: lifting the lower retaining bolt before the entrance will open, three separate lighting actions, the background Spotify volume, four different types of purchase bag, which shelf the Little Mates stock lives on, and setting an inside lock control before you leave so you can't lock yourself out of your own shop.

I froze that list before scoring anything, because it is remarkably easy to move the goalposts once you've seen what came back.

Round one, and I nearly got knocked out

Both of them came back with something better than I wanted them to be. Working pages, embedded players, clickable steps, tick boxes, the lot. For about a minute I sat there thinking the sceptics had it.

Then I started clicking.

ChatGPT: fast, beautiful, and quietly making it up

Sol didn't hang about. Four minutes and forty-four seconds and back it came with a proper interactive web page: embedded player, fifteen clickable timestamps, thumbnails, tick boxes that remembered themselves, "done when" criteria on every task, a verification checklist at the end, and print styling nobody had asked for. It looked like something you'd pay for.

There's a step at 6:45 called "Reset the counter and workstation". Click it and you land in the middle of an explanation of how the women's and men's merchandising layout works. "Make the final merchandising sweep" at 7:35 takes you to the bit where the iPad and card reader go back on charge. "Secure the entrance" at 8:10 drops you into the lights-off sequence. And "Complete the final check and leave" at 8:40 is a tidy generic ending that has quietly lost the inside lock control, which is the single action standing between your closer and a night on the pavement.

Of the seven actions I could match closely enough to score at all, the median error was ten seconds and the worst was over two minutes.

Seven. Out of thirty-eight.

That's the part that took longer to notice and worries me more. Gone: the retaining bolt, the bag rules, the tissue paper, the air conditioning remote, the Little Mates shelf, the lockout control, and eight more besides. In their place it added a pre-entry damage inspection and a generic opening walk-through, neither of which happens in the video.

It skimmed, wrote a confident and professional-looking retail guide out of what it already knew about shops, and handed it over with the air of a consultant who has read the first page of the brief. Unless you have the original video open beside it going line by line, you would read that page and think it was fine.

Claude: stand back, I need thirty-six minutes

Fable 5 Max approached this like a heavyweight who wants you to know he's concentrating. Thirty-six minutes of spinner. Not all of that was the model working, but from where I was sitting it was thirty-six minutes.

And it was worth watching, honestly. What arrived was genuinely handsome: a polished standalone page, twenty-five clickable controls and a valid copy of the video packaged alongside it. Against the 38-action key it had captured 36 of them at some level. It saw the different bags. It saw the tissue paper. It understood an extraordinary amount of what happened in that shop.

Then I clicked the timestamps.

On the actions both systems captured, Claude's median error was seven seconds. Across its full mapped output it was eight. Eight actions were more than twenty seconds out, three were more than thirty, and the worst missed by fifty-one seconds.

The pattern was consistent. Three separate lighting actions sat behind one early timestamp. The paper, zip and canvas bags shared another. The garment bag was bundled in with the tissue paper and the stickers. Everything was in there somewhere, gathered into clumps, and a clump is no use to somebody who needs one specific thing at eight o'clock on a Tuesday.

So one of them barely watched and answered in four minutes, and the other watched almost everything and took half an hour about it, and both come apart at the same inch. Nobody reads a guide from the top. They open it because they're stood at the till at eight o'clock trying to remember which bag the jumper goes in, and they tap the step that matches their problem. Land them ninety seconds away and they shut the guide and shout for a manager, which is the exact event the guide existed to prevent.

I ran it all again, and this is the bit that stopped me

Before publishing any of this I ran both models a second time. Same video, same prompt, nothing changed. I wanted to know whether I'd caught them on a bad day.

ChatGPT got worse. The second run produced fourteen steps, matched four of my 38 actions, and invented an alarm procedure. Not one of its four usable timestamps landed within twenty seconds of the right moment. The score moved. The failure was the same one: big omissions, vague timings, and confident inventions filling the gaps.

Then Claude ran for twenty-six minutes and handed me a guide to a café.

An espresso machine. A food display. An A-board to put out on the pavement. A cash drawer, a shutter, an alarm keypad. There is no café in that video. There is no espresso machine, no A-board and no shutter. It's a clothing shop, and the guide I was reading was a complete, polished, plausible opening and closing procedure for a business that does not exist.

At no point did it tell me anything had gone wrong. No error, no warning, no "I had trouble reading this video". It looked exactly like the good run had looked. Same tidy formatting, same confident tone, same air of a finished product.

I want to be careful here. Forty minutes earlier the same model had understood almost everything in that shop. Both of those results came out of the same model, from the same video and the same prompt, within the hour.

That's the finding. A grounded result and an invented one arrive looking identical, and the only way to tell which one you've got is to sit down with the original video and check all 38 things yourself.

Now picture the owner who did what everyone keeps telling me to do. Uploads the video at nine at night, gets a handsome guide back, skims it, sends it to eleven people. Some of those people are new, most of them will assume the thing in their hand is right, and one of them is going to be stood in a clothing shop at seven in the morning looking for an espresso machine.

The scorecard

Lattify produced nine top-level chapters containing 34 timestamped substeps, and became usable in three minutes and sixteen seconds, which was quicker than either of them. That count of 34 is the shape of its output rather than a score.

In the preliminary mapping, those steps between them represented 34 of the key's 38 actions. On the 33 actions represented in both guides, Lattify's generated start times sat a median of 0.16 seconds from the annotated control against Claude's 7.02, and 21 of Lattify's 33 landed within five seconds against 12 of Claude's. Lattify also pulled the operational furniture out separately: ten tools including the Dyson, the Windex, the air conditioning remote and the lint roller, seven materials including the tissue paper and each type of bag, and a specific warning about the lockout control.

It missed things too. Four of the 38 didn't make it, including the middle display mix and the Little Mates location, and one storage instruction came out vaguer than the video was.

All three guides were scored exactly as generated. No step was added, rewritten or retimed, and no generated timestamp was corrected before it was measured. In practice no guide reaches a member of staff that way. It goes through the person who filmed the job, in an editor where a wrong step gets dragged into place, reworded or deleted in a few seconds by the one person in the building who knows what actually happened. Everything on this list got something wrong, Lattify included. The difference is what you can do about it at nine at night when you spot it.

All five runs, side by side:

RunUsable inActions representedMedian errorWithin 5sInvented content
Lattify3:1634 of 380.16s*21 of 33none noted
Claude Fable 5 Max, run 135:4036 of 387.02s*12 of 33none major
Claude Fable 5 Max, run 226:20not scorablen/an/aan entire café
GPT-5.6 Sol, run 14:447 of 38 scorable10.06s†n/a2 procedures
GPT-5.6 Sol, run 26:174 of 3830.59s†0alarm procedure

* Lattify's and Claude's medians are measured across the 33 actions represented in both guides. † ChatGPT's medians cover only the handful of actions that could be scored at all, so they are not comparable with the other two.

Two caveats, because I would want them if I were reading this. One video, in one kind of shop. And the scoring pass was done with Codex against the frozen key and reviewed by me, so it was neither blinded nor independently adjudicated, which makes this a pilot rather than a formal benchmark.

Why the numbers came out like that

Eight seconds sounds like nothing. It matters because of what a step is for.

If a step is pinned to the exact moment an action starts and stops, the video plays that one action and then pauses while the person does it. Watch, do, next. Eight seconds out and covering three actions, the guide stops being a set of instructions and becomes an index pointing back at the video, which is roughly what a YouTube chapter list already does, for free.

Both of them behaved like something estimating. A model watches a video, forms a view of roughly when something happened, and writes a number down. Most of the time it's close, which is exactly the problem.

Lattify never asks the model to invent a number. The model identifies bounded segments of the transcript, code resolves those segments to source-linked timestamps and throws out any range that doesn't hold up. That doesn't guarantee it spotted every action, and it plainly didn't here, since four of the 38 went missing. What it does is tie every timestamp back to a specific piece of the source instead of a free-form guess, which is why the median sits at a sixth of a second rather than eight. It's also the least demonstrable thing in the product. Nobody has ever been impressed by a demo of a validation layer.

If it were the obvious move, two systems this good wouldn't have both got it wrong in the same direction on the same video.

Now the bit that actually matters

Here's where I put the gloves down, because I think the whole fight is a distraction.

Say they win. Not today, but next spring. Say the next model nails all 38 actions, hits every timestamp to the frame, spots the retaining bolt and the lockout control and the four kinds of bag. It's coming. I'd bet on it.

Congratulations. You have an HTML file.

What do you do with it on Monday morning? Your closer ticked the boxes on the page, but those ticks live in the browser on her phone, so nobody else can see them and they're gone when she clears her history. The Polish lad who started Tuesday needs it in Polish. Somebody's going to have a question at nine on a Saturday and there is nobody inside an HTML file to answer it. Then the supplier changes in September and you do the whole thing again, and again in March.

You can assemble substitutes, of course. A chat tool to write it, a WhatsApp group to hand it out, a spreadsheet to track who's done it, a phone to translate it, a clipboard for the checklist, and a shared folder where the current version is whichever file has the latest date on it. That's six tools with you as the glue, re-syncing every one of them by hand each time a procedure changes. And if the answer is that a developer could just build the rest, they could build a version of it, and then inherit everything that comes after launch, right down to holding staff data they're now responsible for keeping safe. You've hired yourself as an unpaid IT department for a product with one user.

Delivery is its own problem too, and one the models don't control. Since the fifteenth of January this year, Meta's WhatsApp Business terms have barred general-purpose AI assistants from the platform, while AI running as part of a defined business service is still permitted. Platform rules change, so I wouldn't build a company on that one. It's a fair illustration of the difference though: getting onto the phone in your closer's hand means being a specific business's tool rather than a general one.

When a generated guide is fine

I should be straight about when none of this matters. If the job is short and forgiving, if somebody only needs a rough orientation, if getting it wrong costs very little, or if you're going to be stood next to them anyway, then a generated page is fine. So is a video with chapters on it, or a laminated sheet by the till. Plenty of businesses need nothing more than that and I'd rather say so.

It changes when you want the job done the same way without you there.

What I built

Lattify is the part that happens after the file.

Your best person films the job once. The guide comes back with source-linked timestamps the creator can inspect and correct, the tools and materials pulled out per step, the safety warnings kept where they belong. You check it, fix anything the camera didn't make obvious, and publish. You assign it to the three people who need it, in whichever of ten languages they read. They work through it on their phone, hands-free if their hands are full, and when they get stuck they ask and get an answer out of your guides rather than out of the internet. You see who finished it, where people stalled, and what they asked that nobody had an answer for. Supplier changes in September and everyone has the new version that afternoon.

In this test the guide was usable in three minutes, with the deeper analysis carrying on in the background, because an owner stood in a shop at nine at night doesn't want to wait for a confusion report before they can hand something to their staff.

The shell is getting cheap

The genuinely interesting result of this fight was how good the losing corner looked. Two years ago "upload a video, get an interactive guide back" was a product. Now it's one prompt and four minutes, and it will keep getting better and cheaper.

What stays hard is everything the shell doesn't touch: knowing which bag the jumper goes in, landing on the right second when somebody taps a step, getting onto that worker's phone in her language, telling the owner that four people this week got stuck in the same place.

The video went in and a plausible page came out. The shop's actual knowledge stayed in the video.

If any of this sounded familiar, we built Lattify for exactly this problem.

Join the Waitlist