2026.08.12 (Wed)
2026.08.14 (Fri) updated

โœจ GPT-5.6 Solโ€™s Summary ใ€€

I opened Codexโ€™s result expecting something solid after 16 hours of tests and audits, and it was absolute trash. That one screen started exposing every problem I had missed.

I Let It Run for 16 Hours, and the Result Was Absolute Trash

Yesterday, I spent a long time debating and writing detailed page-by-page specifications for a major UI/UX overhaul. Then I set a Goal telling Codex to build from those specifications.

Codex reporting completion and passing tests after 16 hours and 44 minutes of work

The Goal ran for 16 hours and 44 minutes. I opened it today full of expectation, and the result was absolute trash.

The actual result screen with company and personal identifiers redacted

I was humbled. Again.

When I think about how much I could achieve if I learned to run AI well, I get completely carried away. Then it forgets what I meant, does whatever it wants, and sprays out shit that only looks vaguely convincing. The crash afterward is brutal. This time I had even written a long specification and received a report saying it had passed 16 hours of tests and audits. What I wanted still was not there.

One Failure Exposed Everything Else I Had Missed

As I argued with AI about why this had happened, I learned that several problems I had been wrestling with while using Codex already had names and established methods. People keep running into the same issues, so of course the discussion was already active. I simply had not known because I had shut my ears and paid no attention.

I kept pressing on that one screen, and the problems I had missed started falling out one after another. Instead of forcing them into one post, I recorded each problem I actually retried and the result separately.

That made me think someone must already have turned this into a service, and of course someone had. Find a recurring irritation, scratch it properly, and rake in the money. It sounds so easy when said out loud. Building that cycle and keeping it running is anything but easy.

I still believe that running AI well can produce enormous results. But running it longer, spending more tokens, and passing more tests were not evidence that it understood what I wanted. Before I accept another 16-hour pile of trash, I need to check at the start whether it actually understands the result I am asking for.

Categories: ,

Updated:

Leave a comment