Hesham HassanFrontend Architect
← Writing

Architecting an AI Assistant into a Zero-Build Website

ChefAI is a conversational assistant for professional chefs, built for Unilever Food Solutions. It generates recipes, analyses menus and tailors its advice to the business asking. It runs in several markets and languages, including right-to-left ones. I led the frontend architecture for about four months, with a team that changed shape a few times along the way.

The awkward part was the platform. The site ships on Adobe Edge Delivery Services, which serves files straight out of the repository. No bundler, no transpile step, no framework runtime on the page. What you commit is what the browser runs. That is great for Core Web Vitals, and inconvenient when the feature you have been asked to build is an agentic chat client that wants long-lived connections, streaming state and a proper module graph.

This post is the list of decisions that got us there without adding a compile step. I have tried to write down what each one cost as well as what it bought, because the cost is usually the part that gets left out.

4months as lead architect
0 KBframework JS on content pages
0build steps added
3 minlongest agent run we had to hide
The launch post from Unilever Food Solutions. Worth thirty seconds before the diagrams, so you know what the boxes are describing.

The constraints

There were three, and none of them had any give in it.

No build stepplatform constraint
Core Web Vitals budgetbusiness constraint
Agentic, streaming UXproduct constraint
ChefAIeverything below follows from these three
Most of what follows is what happens when you take all three seriously at the same time instead of quietly dropping one.

How the system ended up

Before the individual decisions, the overall shape. Three entry points share one widget core. The core talks to the agent platform over two channels that carry different things and live for different lengths of time. Identity, local state and configuration sit next to the core rather than inside it.

SURFACEInline assistantauthored into any page
SURFACEFloating modalavailable site-wide
SURFACEGuided onboardingbusiness profiling flow
WIDGET CORE — loaded on demand
ui · presentationhooks · behaviourmodel · canonical message
CHANNEL 1 · REQUESTThe answerone payload, up to 3 min, cumulative
CHANNEL 2 · STREAMThe reasoningmany events, live, ephemeral
AGENT PLATFORM
chatthreadsusersbusinessrecommendations
IdentityAnonymous from the first message, merged into the account on login
Local stateConversation pointer plus a cache that renders before the network answers
ConfigurationEndpoints and keys authored as content, resolved per environment
The two channels are the important bit. Almost every decision below comes from keeping them apart rather than treating them as one response.

1. React as a feature's runtime, not the site's dependency

The chatbot needed component-shaped UI. The site was not allowed to ship a framework. Neither side was going to move, so I stopped thinking of React as a dependency of the site and started thinking of it as something one feature loads for itself.

React comes from a CDN, the first time a chat surface actually mounts. A page where nobody opens the assistant never downloads it. Module aliasing comes from an import map declared once in the document head. That gives us most of what a bundler's alias config gives you, with no tooling behind it.

Context

We needed rich UI, and the content pages were not allowed to pay for it.

Decision

Load the view library at runtime, scoped to the feature. Use import maps for module aliasing instead of a bundler.

Consequence

Content pages ship no framework. Module resolution moved to runtime, so static analysis got weaker and we had to switch off one lint rule on purpose.

Inside the core, two rules did most of the work. Dependencies only point downward, so presentation never reaches for transport and transport never knows what a message bubble looks like. And a file that only re-exports other files is not allowed to exist. Wrapper modules are how a clean layering turns into a maze within six months. Both rules can be checked in a thirty-second review, which is the only reason they survived four months and a dozen contributors.

2. The client mints the correlation ID

If I could only keep one decision from this project, it would be this one.

The agent backend gives you two things: a request that eventually returns the final answer, and a separate event stream with the agent's intermediate reasoning. The obvious wiring is to send the message, get a run ID back, then subscribe to that run's stream. It is also broken, quietly. By the time you have the ID, the first reasoning events have already been emitted to nobody.

So we flipped it. The client generates the run ID, opens the stream, and only then sends the message with that ID attached.

▲ SERVER OWNS THE ID — the listener arrives late

send message→server mints ID→response returns→subscribe→✗ early reasoning lost

▼ CLIENT OWNS THE ID — the listener exists before the work

mint ID→open stream→send message with that ID→✓ nothing missed
Same amount of code, different order. The ID becomes an input to the request instead of an output of it, and the race simply is not there any more.

It cost nothing. No replay buffer on the server, no reconnect-and-catch-up logic, no sleeping for 200ms and hoping. When the recommendations engine came along later it reused the same pattern without changes.

The transport choice came from the same place. The textbook client for server-sent events is EventSource, and we did not use it.

EventSource
  • Cannot send custom headers, and every call here is authenticated
  • No cancellation tied to a component lifecycle
  • Reconnects automatically, on its own schedule
Streamed fetch
  • Full control of headers and auth
  • Abort signal wired to unmount, so closing the panel closes the stream
  • We decide what "done" means
The price is that we parse SSE frames ourselves, which EventSource would have done for us. Auth was not optional, and a stream that outlives its component is a leak waiting for a bug report, so I still think it was the right call.

3. The stream is not the answer

The event stream does not carry the response. It carries the agent's reasoning: "checking your menu", "finding seasonal dishes", that sort of thing. The actual answer arrives separately. If you concatenate them into one text field, users see the model's scratchpad presented as professional advice, which is not a good look for a product aimed at chefs.

So a pending message has two channels with two lifetimes. Reasoning text is ephemeral. Each event replaces the previous one and none of it makes it into the transcript. Response text accumulates, and that is what gets persisted.

REASONING · replaced each event · never kept
Looking at your menu…
RESPONSE · accumulated
For a spring menu, I'd start with…
FINAL · recipes, products, follow-ups
Full answer with structured content attached
One message, three states. Someone scrolling back through yesterday's conversation sees the advice, not the process that produced it.

The same transport later served a second surface where the opposite was true. In the personalised onboarding flow, the reasoning is the content: it narrates progress while the agent builds a business profile. Same stream, two contracts, one client. That only worked because "what does this stream mean" was a parameter, not something baked into the UI.

4. Perceived latency is an architecture problem

Agent runs can take up to three minutes, and the answer lands as one payload at the end.

A block of text that appears all at once feels slower than text that arrives progressively, even when the wait is identical. So once the response is in, the interface re-streams it word by word, picking up from wherever the live reasoning stopped.

For a spring menu I'd lead with charred asparagus, then…
This is presentation, not transport, and I want to be clear about that. The seam is already the right shape for real token streaming if the backend ever offers it.

I am comfortable with this. It is not pretending to a capability we do not have. It is matching the interface to how people actually read. The alternative was a three-minute spinner, which is a bigger lie.

5. Degrade in tiers

Long-running AI calls fail in more ways than a CRUD request does. Rather than let each failure invent its own behaviour, we wrote the ladder down.

T1Full experience. Live reasoning, then the streamed answer.
T2Stream fails. Abort it and fall back to the plain request. No narration, full answer. Most users never notice.
T3Request fails. The conversation survives. Only the pending message becomes something you can retry.
A dead stream never costs you the answer. A dead request never costs you the conversation.

Two smaller choices hold this up. Every long request races an explicit timeout, so a hung agent turns into a real error instead of a spinner that never resolves. And thread resolution repairs itself rather than assuming the best: a stale thread gets replaced, a missing user gets recreated and the call retried once. Backend state drifts. Deployments happen, tokens expire, someone wipes a test environment. A client that assumes otherwise ends up showing users a dead end for somebody else's operational event.

6. Identity accrues instead of gating

Asking a chef to register before asking their first question would have killed the funnel. So identity builds up over time.

The merge is the hard part. The detail that bit us was making sure it survives the redirect that follows login. Get that wrong and you orphan the exact conversation that earned you the signup.

Conversation history follows the same idea. It renders immediately from a local cache and revalidates against the network in the background. The cache is keyed by conversation, and that is not optional: a cache that shows the previous conversation for even one frame is not a performance bug, it is a privacy incident.

7. Configuration is content

On a platform with no build step, environment configuration has no business being in code. Endpoints and keys are authored in a table with one column per environment, and resolved when the module loads.

Repointing an environment is a content publish. No rebuild, no redeploy, no engineer needed. That mattered more than it sounds, because it took a recurring release-day dependency off our plate and gave it to the delivery team, who were better placed to handle it anyway.

The performance budget got the same treatment. The assistant is heavy: a view library, a markdown pipeline, a sanitiser, a handful of modules and stylesheets. None of it is on the critical path. It gets warmed during idle time on the one page where the next click is predictable.

critical path
deferred
idle
prefetch assistant
on click
The expensive work happens while nobody is waiting. Opening the assistant feels instant, and a visitor who never opens it pays nothing for it.

8. A streaming surface still has to work for everyone

Streaming text and floating panels are two of the easiest places to build something only sighted mouse users can operate, without noticing you have done it. We treated accessibility as a fourth constraint next to the three at the top, not a pass at the end.

Context

A chat surface that streams reasoning and opens as a floating modal is exactly the shape that breaks first for keyboard and screen reader users.

Decision

Make accessibility an input to the same decisions that shaped everything else (the stream, the modal, the motion) rather than a separate audit afterwards.

Consequence

Each feature carries its accessible behaviour in the same module as its visual behaviour, so the two cannot drift apart in a later refactor.

Three concrete commitments. Each one maps to something a chef using a screen reader or a keyboard would otherwise have hit on day one.

Right-to-left got the same treatment instead of a CSS toggle bolted on at the end. Direction is a property of the content, so a thread that mixes English and Arabic keeps each language ordered correctly without leaking direction into the surrounding chrome. Once you have done both, accessibility and internationalisation start to look like the same discipline. Both come down to not assuming the reader is you.

What it cost

Every one of these decisions has a bill attached. Here is ours.

We own a stream parser. Skipping EventSource bought us auth and lifecycle control, and added protocol handling to our maintenance surface for good.
Progressive text is perception. Not real token streaming. Right for the contract we had. I would want it revisited the moment that contract changes.
Runtime module resolution. Import maps cost us static verification of imports. We knew that going in.
Single-conversation cache. Simple and bounded, and wrong the day we ship parallel threads. The ceiling is documented.
Throttled live regions need tuning. Too eager and the region spams. Too sparse and updates feel broken. The interval came out of user testing, not a guess.

The part that was leadership rather than architecture

Most of these decisions were cheap to make and expensive to keep. Under delivery pressure, the pull is always towards "just add a bundler", "just put the token where it is easy to reach", "just let this component call the API directly this once".

Three things kept the design intact across four months and a dozen contributors. We wrote the layering rules down next to the code, so a review had something to point at instead of an opinion to defend. We made the boundaries cheap to check. No wrapper modules, dependencies point downward, transport stays in one layer: all things you can verify in seconds. And when a rule got broken for a good reason, we recorded why, so the next person inherited a decision rather than a mystery.

If the architecture only exists in the architect's head, it is a preference, not an architecture.

What I would take to the next project

Own your correlation IDs. Whenever subscribing and starting the work are separate calls, let the client pick the identity. It turns a race condition into a non-event for the price of one function.

Put the ugliness where it can be deleted. Backend contracts drift. One deliberately defensive normalisation boundary kept every component clean and made every rename a one-file change. Spread that same tolerance across twelve files and you have a problem you cannot find.

Streaming is a UX contract before it is a transport one. The hard question was never how to read from a socket. It was noticing that an agent's reasoning and an agent's answer are different kinds of text that deserve different lifetimes on screen. Get that wrong and no amount of transport correctness saves the experience.