Skip to content

Eight Weeks of AI R&D

Pixel art terminal window titled week 8 — Eight Weeks of AI R&D — with two blue geese and a white robot goose talking at a lamplit desk beneath an R&D roadmap corkboard, mugs labelled Image AI and Website AI between them

What eight weeks of AI R&D in a real agency actually looks like

The honest answer is that it looks a lot less linear than you’d expect, and a lot more useful than a personal project ever could be.

Two tools. Two builders. Eight weeks. And enough shared ground between the projects that we’ve spent as much time learning from each other as we have from the work itself.

Two tools that do very different things — and what that’s worth

Image AI and Website AI are not the same problem. One is about generating and editing visual assets within tight brand constraints. The other is about turning a structured brief into a fully functioning website. The models that work well for one don’t necessarily work well for the other. The user experience considerations are different. The failure modes are different.

But that difference has been one of the most useful things about the co-op structure. Because we’ve both been watching each other build something we’re not responsible for, we can see things the other person can’t. When you’re deep in your own project, certain decisions stop looking like decisions — they just look like the way things are. Stepping into someone else’s tool for feedback breaks that. You notice what’s confusing immediately, because you don’t have the context that made it feel obvious to the person who built it.

It can be hard to context switch. Going from one project’s logic to another mid-week takes more energy than staying in one lane. But the feedback loop it creates is worth it. The blind spots we’ve caught in each other’s work are the kind that internal testing almost never surfaces.

There’s also the shared infrastructure layer — the brand intelligence schema, the data model that both tools pull from. Decisions about how that gets structured affect both pipelines. That coordination has required a level of communication that a solo project never would, and the tools are better for it.

What a real agency makes possible

A personal project gives you control. It doesn’t give you a designer sitting next to you telling you whether the output is actually what they’d use.

Working inside a real agency means the people who would use these tools are accessible. Designers, social teams, account leads — people with opinions about what good looks like that come from doing the work, not from evaluating a demo. Feature ideas that come from those conversations are different from feature ideas that come from us. They reflect actual workflow problems, actual friction points, actual things that would make someone’s day easier or not.

Some of the creative assets being generated through Image AI right now are going out for real client approval. That’s a different standard than a proof of concept. It means the quality bar is set by someone who has no investment in the tool working — only in the output being good enough to use. That pressure is useful in a way that’s hard to replicate outside of a real working environment.

Where our heads are at

Week one had a lot of hypotheticals. Schemas that existed as documents, pipelines that existed as diagrams, tools that existed as plans. Eight weeks in, those things are real. The brand intelligence schema isn’t a concept anymore — it’s a structure both tools depend on. The pipelines aren’t diagrams — they’re running.

That shift changes how the work feels. It’s less abstract. The decisions have weight because the thing exists and people are using it. The end result isn’t hypothetical anymore — it’s visible, testable, and getting closer every week.

There’s also a lot less confusion about direction. The first few weeks involved a lot of figuring out what we were actually building. Now both projects have enough shape that the remaining work is clear. That clarity makes the second half feel different from the first.

Working with AI to build AI tools

The thing that’s become obvious over eight weeks is that AI is most useful when you’re working alongside it, not handing things off to it. The best results — in the tools and in the build process itself — have come from step by step collaboration. Prompt, review, redirect, build. Not: here’s the full plan, come back in three days.

That applies inside the tools too. The AI is there to fill in the gaps the user leaves — but it has to be trained well enough to fill them in the right direction when inputs are vague. That takes a lot of testing and a lot of iteration. A human eye at every stage is what keeps the output from drifting.

We’ve used AI to help build, to help document, to help generate inside the tools themselves. It’s a very capable collaborator when you treat it like one — meaning you stay in the loop, you check the output, and you don’t mistake fluency for correctness.

Done enough, not finished

Both tools are in a good place. The core features work. The outputs are getting better with every round of testing. But there’s a version of “done” that means functional and a version that means polished, integrated, and ready for production — and those are different things.

Integration will surface new things that need changing. More client testing will too. That’s not a problem — that’s just what the process looks like when you’re building something real. By the end of the co-op both tools will be functioning and polished. And after that, there will still be room to optimise, especially as the models themselves keep improving.

That’s probably the most honest thing about building with AI right now. The tools get better as the technology gets better. The work doesn’t really end — it just reaches a version worth shipping.