Architecting an AI Assistant into a Zero-Build Website
ChefAI is a conversational assistant for professional chefs, built for Unilever Food Solutions. It generates recipes, analyses menus and tailors its advice to the business asking. It runs in several markets and languages, including right-to-left ones. I led the frontend architecture for about four months, with a team that changed shape a few times along the way.
The awkward part was the platform. The site ships on Adobe Edge Delivery Services, which serves files straight out of the repository. No bundler, no transpile step, no framework runtime on the page. What you commit is what the browser runs. That is great for Core Web Vitals, and inconvenient when the feature you have been asked to build is an agentic chat client that wants long-lived connections, streaming state and a proper module graph.
This post is the list of decisions that got us there without adding a compile step. I have tried to write down what each one cost as well as what it bought, because the cost is usually the part that gets left out.
The constraints
There were three, and none of them had any give in it.
How the system ended up
Before the individual decisions, the overall shape. Three entry points share one widget core. The core talks to the agent platform over two channels that carry different things and live for different lengths of time. Identity, local state and configuration sit next to the core rather than inside it.
1. React as a feature's runtime, not the site's dependency
The chatbot needed component-shaped UI. The site was not allowed to ship a framework. Neither side was going to move, so I stopped thinking of React as a dependency of the site and started thinking of it as something one feature loads for itself.
React comes from a CDN, the first time a chat surface actually mounts. A page where nobody opens the assistant never downloads it. Module aliasing comes from an import map declared once in the document head. That gives us most of what a bundler's alias config gives you, with no tooling behind it.
Context
We needed rich UI, and the content pages were not allowed to pay for it.
Decision
Load the view library at runtime, scoped to the feature. Use import maps for module aliasing instead of a bundler.
Consequence
Content pages ship no framework. Module resolution moved to runtime, so static analysis got weaker and we had to switch off one lint rule on purpose.
Inside the core, two rules did most of the work. Dependencies only point downward, so presentation never reaches for transport and transport never knows what a message bubble looks like. And a file that only re-exports other files is not allowed to exist. Wrapper modules are how a clean layering turns into a maze within six months. Both rules can be checked in a thirty-second review, which is the only reason they survived four months and a dozen contributors.
2. The client mints the correlation ID
If I could only keep one decision from this project, it would be this one.
The agent backend gives you two things: a request that eventually returns the final answer, and a separate event stream with the agent's intermediate reasoning. The obvious wiring is to send the message, get a run ID back, then subscribe to that run's stream. It is also broken, quietly. By the time you have the ID, the first reasoning events have already been emitted to nobody.
So we flipped it. The client generates the run ID, opens the stream, and only then sends the message with that ID attached.
▲ SERVER OWNS THE ID — the listener arrives late
▼ CLIENT OWNS THE ID — the listener exists before the work
It cost nothing. No replay buffer on the server, no reconnect-and-catch-up logic, no sleeping for 200ms and hoping. When the recommendations engine came along later it reused the same pattern without changes.
The transport choice came from the same place. The textbook client for server-sent events is EventSource, and we did not use it.
EventSource
- Cannot send custom headers, and every call here is authenticated
- No cancellation tied to a component lifecycle
- Reconnects automatically, on its own schedule
Streamed fetch
- Full control of headers and auth
- Abort signal wired to unmount, so closing the panel closes the stream
- We decide what "done" means
3. The stream is not the answer
The event stream does not carry the response. It carries the agent's reasoning: "checking your menu", "finding seasonal dishes", that sort of thing. The actual answer arrives separately. If you concatenate them into one text field, users see the model's scratchpad presented as professional advice, which is not a good look for a product aimed at chefs.
So a pending message has two channels with two lifetimes. Reasoning text is ephemeral. Each event replaces the previous one and none of it makes it into the transcript. Response text accumulates, and that is what gets persisted.
Looking at your menu…
For a spring menu, I'd start with…
Full answer with structured content attached
The same transport later served a second surface where the opposite was true. In the personalised onboarding flow, the reasoning is the content: it narrates progress while the agent builds a business profile. Same stream, two contracts, one client. That only worked because "what does this stream mean" was a parameter, not something baked into the UI.
4. Perceived latency is an architecture problem
Agent runs can take up to three minutes, and the answer lands as one payload at the end.
A block of text that appears all at once feels slower than text that arrives progressively, even when the wait is identical. So once the response is in, the interface re-streams it word by word, picking up from wherever the live reasoning stopped.
I am comfortable with this. It is not pretending to a capability we do not have. It is matching the interface to how people actually read. The alternative was a three-minute spinner, which is a bigger lie.
5. Degrade in tiers
Long-running AI calls fail in more ways than a CRUD request does. Rather than let each failure invent its own behaviour, we wrote the ladder down.
Two smaller choices hold this up. Every long request races an explicit timeout, so a hung agent turns into a real error instead of a spinner that never resolves. And thread resolution repairs itself rather than assuming the best: a stale thread gets replaced, a missing user gets recreated and the call retried once. Backend state drifts. Deployments happen, tokens expire, someone wipes a test environment. A client that assumes otherwise ends up showing users a dead end for somebody else's operational event.
6. Identity accrues instead of gating
Asking a chef to register before asking their first question would have killed the funnel. So identity builds up over time.
Conversation history follows the same idea. It renders immediately from a local cache and revalidates against the network in the background. The cache is keyed by conversation, and that is not optional: a cache that shows the previous conversation for even one frame is not a performance bug, it is a privacy incident.
7. Configuration is content
On a platform with no build step, environment configuration has no business being in code. Endpoints and keys are authored in a table with one column per environment, and resolved when the module loads.
Repointing an environment is a content publish. No rebuild, no redeploy, no engineer needed. That mattered more than it sounds, because it took a recurring release-day dependency off our plate and gave it to the delivery team, who were better placed to handle it anyway.
The performance budget got the same treatment. The assistant is heavy: a view library, a markdown pipeline, a sanitiser, a handful of modules and stylesheets. None of it is on the critical path. It gets warmed during idle time on the one page where the next click is predictable.
8. A streaming surface still has to work for everyone
Streaming text and floating panels are two of the easiest places to build something only sighted mouse users can operate, without noticing you have done it. We treated accessibility as a fourth constraint next to the three at the top, not a pass at the end.
Context
A chat surface that streams reasoning and opens as a floating modal is exactly the shape that breaks first for keyboard and screen reader users.
Decision
Make accessibility an input to the same decisions that shaped everything else (the stream, the modal, the motion) rather than a separate audit afterwards.
Consequence
Each feature carries its accessible behaviour in the same module as its visual behaviour, so the two cannot drift apart in a later refactor.
aria-live="polite"throttled, not per token. A screen reader narrating every word is unusable, whatever the audit saysprefers-reduced-motionthe diagrams in this post follow the same rule as the productRight-to-left got the same treatment instead of a CSS toggle bolted on at the end. Direction is a property of the content, so a thread that mixes English and Arabic keeps each language ordered correctly without leaking direction into the surrounding chrome. Once you have done both, accessibility and internationalisation start to look like the same discipline. Both come down to not assuming the reader is you.
What it cost
Every one of these decisions has a bill attached. Here is ours.
The part that was leadership rather than architecture
Most of these decisions were cheap to make and expensive to keep. Under delivery pressure, the pull is always towards "just add a bundler", "just put the token where it is easy to reach", "just let this component call the API directly this once".
Three things kept the design intact across four months and a dozen contributors. We wrote the layering rules down next to the code, so a review had something to point at instead of an opinion to defend. We made the boundaries cheap to check. No wrapper modules, dependencies point downward, transport stays in one layer: all things you can verify in seconds. And when a rule got broken for a good reason, we recorded why, so the next person inherited a decision rather than a mystery.
If the architecture only exists in the architect's head, it is a preference, not an architecture.
What I would take to the next project
Own your correlation IDs. Whenever subscribing and starting the work are separate calls, let the client pick the identity. It turns a race condition into a non-event for the price of one function.
Put the ugliness where it can be deleted. Backend contracts drift. One deliberately defensive normalisation boundary kept every component clean and made every rename a one-file change. Spread that same tolerance across twelve files and you have a problem you cannot find.
Streaming is a UX contract before it is a transport one. The hard question was never how to read from a socket. It was noticing that an agent's reasoning and an agent's answer are different kinds of text that deserve different lifetimes on screen. Get that wrong and no amount of transport correctness saves the experience.