Skip to content
Fantástico Mundo de Jon
RSS

A website from scratch via chat: HTML is ready, CSS isn't

I built an entire website by copying and pasting from the chat. The structure was almost ready, but the presentation wasn't — because HTML has a correct answer and layout only has a correct answer for a context that the model doesn't know.

In a hurry? Ask Claude for the TL;DR — it reads the page and summarises it.

en Machine translation by mistral-small3.2:24b, reviewed by the author. Read the Portuguese original

The markup of an entire website comes out of the chat ready to paste. The stylesheet doesn’t come out—and the reason isn’t that the model is worse at CSS than it is at HTML.

Over the past few weeks, I built a small website from start to finish by chatting in the browser: cover page, three internal pages, contact form. No agent, no plugin in the editor, nothing that wrote files for me. I asked, it responded, I selected, Ctrl+C, Ctrl+V in VS Code, saved, refreshed the tab. The dumbest method available today, and on purpose: there’s Cursor, there’s Copilot, there’s Vercel’s v0 generating components with preview—I specifically wanted the scenario without a tool holding things together because that’s how most people actually use it.

The result has an asymmetry that I didn’t expect to be so clean. The structure came out almost ready. The presentation didn’t come out at all.

How It Was Made

Period
nov–dec 2024
Models
GPT-4o and Claude 3.5 Sonnet
Method
copy and paste, no agent
Editor
VS Code, no AI plugin
Stack
HTML + Tailwind 3
Cost
US$ 40/month (ChatGPT Plus + Claude Pro, US$ 20 each)

Markup has a template; layout has context

This is the entire sentence of the post, and it’s worth spending a paragraph on it before the examples.

For a given content, there is a correct markup. An email field is input type="email" with a label associated by for/id. This doesn’t depend on the client, the brand, the audience, or my taste. It’s a public convention, written in specification, repeated on millions of pages—exactly the kind of thing that a language model learns well because the answer is the same regardless of who asks.

Layout doesn’t work like this. What is the correct space between the section title and the first paragraph? There is no answer outside of context: it depends on the size of the body text, the width of the column, the density of the rest of the page, how much content will actually go there, whether the site is being read on the subway or a 27“ monitor. None of this information is in my prompt. The model doesn’t get the answer wrong—it returns the median of all plausible answers, which is a correct answer for an average project that isn’t mine.

It’s the difference between a test question and a project question. Chat is excellent at the first and structurally limited in the second.

What comes out ready really is ready

Being fair to the part that works, because it’s the largest part of the code volume.

I asked for the contact form in one sentence—name, email, phone, message—and this came back:

<form action="/contato" method="post" novalidate>
  <div>
    <label for="nome">Nome</label>
    <input id="nome" name="nome" type="text" autocomplete="name" required />
  </div>
  <div>
    <label for="email">E-mail</label>
    <input id="email" name="email" type="email"
           autocomplete="email" inputmode="email" required />
  </div>
  <div>
    <label for="telefone">Telefone</label>
    <input id="telefone" name="telefone" type="tel"
           autocomplete="tel" inputmode="tel" />
  </div>
  <div>
    <label for="mensagem">Mensagem</label>
    <textarea id="mensagem" name="mensagem" rows="6" required></textarea>
  </div>
  <button type="submit">Enviar mensagem</button>
</form>

Look at what’s there without me asking: correct autocomplete in each field, inputmode for mobile keyboard, type="tel" instead of text, label tied by for, required only on the fields that make sense, the button with explicit type="submit". This is better than many forms I’ve reviewed in production written by people.

The same goes for obvious componentization. Ask for a header, a footer, and a service card, and the cut between them comes in the place any dev would do. And the placeholder text is decent: instead of lorem ipsum, it says “Service within 24 business hours” and “Quote without obligation,” which is bad as final copy and great as scaffolding—you can diagram on top of it and see the layout with real-sized text, in Portuguese, with accents, where lorem ipsum deceives.

None of this is insignificant. It’s the boring part of the job, done in seconds, with a high success rate.

The markup is correct and the semantics are average

Now the first crack, which is subtle because it doesn’t look like a crack: valid HTML isn’t the same as well-structured HTML.

What came for the cover:

<div class="header">
  <div class="logo">Ateliê Marena</div>
  <div class="menu">
    <div class="menu-item"><a href="/servicos">Serviços</a></div>
    <div class="menu-item"><a href="/contato">Contato</a></div>
  </div>
</div>

<div class="hero">
  <div class="hero-title">Móveis sob medida</div>
  <div class="hero-text">Projeto, marcenaria e instalação.</div>
</div>

<div class="services">
  <div class="service-card">
    <div class="service-title">Projeto 3D</div>
    <div class="service-desc">Você vê antes de aprovar.</div>
  </div>
</div>

Renders perfectly. Passes the validator. And it’s a page with no reference points for someone who can’t see the screen: no landmarks, no heading levels, nothing that a screen reader can use to jump directly to the content or list the sections of the page.

What I had to write by hand:

<header>
  <a href="/" class="logo">Ateliê Marena</a>
  <nav aria-label="Principal">
    <ul>
      <li><a href="/servicos">Serviços</a></li>
      <li><a href="/contato">Contato</a></li>
    </ul>
  </nav>
</header>

<main>
  <section aria-labelledby="titulo-hero">
    <h1 id="titulo-hero">Móveis sob medida</h1>
    <p>Projeto, marcenaria e instalação.</p>
  </section>

  <section aria-labelledby="titulo-servicos">
    <h2 id="titulo-servicos">Serviços</h2>
    <article>
      <h3>Projeto 3D</h3>
      <p>Você vê antes de aprovar.</p>
    </article>
  </section>
</main>

The verdict is the same in all cases I saw: the model chooses div when the right element exists because div is the answer that is never wrong. It’s not ignorance—if I ask “what element should I use for navigation?”, it answers nav without hesitation and explains it perfectly. It just doesn’t choose the right element when no one demands it, just like no one demanded it in most of the HTML on the internet where it was trained.

The contrast fails without anyone noticing

This is my favorite because it’s the error that survives all review stages a common person does.

I asked for “a highlight button, with the site’s main color.” It returned:

<button class="bg-indigo-500 text-white px-6 py-3 rounded-lg">
  Solicitar orçamento
</button>

<p class="text-gray-400 mt-2">Resposta em até 24 horas úteis.</p>

It looks nice. It looks exactly like what you imagined when reading the paragraph. And both lines fail the WCAG contrast criterion.

White on #6366f1—Tailwind 3’s indigo-500—gives 4.47:1. The minimum for normal text in AA is 4.5:1. It fails by three hundredths, a margin that no human eye detects and no browser warns about. The text-gray-400 is worse and more common: #9ca3af on white gives 2.54:1, almost half of what’s required, and it’s the default class for “secondary text” in practically every component that comes out of the chat.

Color pair Contrast (n:1) AA normal text
#ffffff on #6366f1 4.47 fail
#9ca3af on #ffffff 2.54 fail
#6b7280 on #ffffff 4.83 pass
#ffffff on #4f46e5 6.29 pass
Calculation by the WCAG 2.x relative luminance formula, not measurement—two hexadecimals are enough. The first two rows are what the chat returned; the last two are the correction, one step in the Tailwind scale. AA requires 4.5:1 for normal text and 3:1 for large text.

The calculation fits in ten lines and doesn’t depend on any library:

const canal = (v) => (v <= 0.03928 ? v / 12.92 : Math.pow((v + 0.055) / 1.055, 2.4));

const luminancia = (hex) => {
  const [r, g, b] = [1, 3, 5].map((i) => parseInt(hex.slice(i, i + 2), 16) / 255);
  return 0.2126 * canal(r) + 0.7152 * canal(g) + 0.0722 * canal(b);
};

const contraste = (a, b) => {
  const [x, y] = [luminancia(a), luminancia(b)].sort((m, n) => n - m);
  return (x + 0.05) / (y + 0.05);
};

#ffffff on #6366f1 4.47:1 fail #9ca3af on #ffffff 2.54:1 fail #6b7280 on #ffffff 4.83:1 pass #ffffff on #4f46e5 6.29:1 pass

saída

The correction is ridiculously cheap: indigo-500 becomes indigo-600 and white gets 6.29:1; gray-400 becomes gray-500 and the secondary goes up to 4.83:1. One digit in the hexadecimal separates failing from passing. The problem was never the difficulty of the correction—it was no one knowing there was something to correct.

The class that doesn’t exist doesn’t cause an error

At some point this appeared in my HTML:

<div class="flex-center gap-4 text-md shadow-soft hover:scale-102">

Four of these six classes don’t exist in Tailwind 3. gap-4 exists. hover: is a valid prefix. flex-center, text-md, shadow-soft, and scale-102 are inventions—plausible utilities, with the names they should have, that anyone would swear are part of the framework.

The cruel detail: text-md doesn’t exist because Tailwind’s scale is text-sm, text-base, text-lg. scale-102 doesn’t exist because the scale jumps from scale-100 to scale-105. flex-center is the obvious fusion of flex items-center justify-center, which everyone has already wanted to exist. They are exactly the names a human would guess.

The same happened with Bootstrap 5 when I tested the same page on the other framework:

<div class="d-flex-center mt-6 text-medium">

mt-6 doesn’t exist: Bootstrap 5’s spacing scale goes from 0 to 5. d-flex-center and text-medium don’t either. And—this is the point—none of this emits an error. There’s no warning in the console, no red underline in the editor, the build doesn’t fail. The class simply generates no rule. The element remains unstyled, you look, think the spacing “got a bit tight,” and correct the symptom on the neighboring element.

I found all the invented classes of the project at once, and none of them by looking at the screen. I discovered all of them at once, looking at the generated CSS and searching for what had no output.

Every hover without focus erases keyboard navigation

This error is 100% consistent. Every time I asked for an interaction state, only the mouse one came back:

.botao {
  background: #4f46e5;
  color: #ffffff;
  transition: background 150ms;
}

.botao:hover {
  background: #4338ca;
}

It works well for those who use a mouse and doesn’t exist for those who use a keyboard. Worse: if at some point the CSS gets an outline: none—and it does, because “removing that ugly blue border from Chrome” is a common request and the model complies without discussing—then the person navigating by Tab completely loses track of where they are on the page.

What the block needs to have:

.botao:hover {
  background: #4338ca;
}

/* o par que nunca vem junto */
.botao:focus-visible {
  outline: 2px solid #b23c17;
  outline-offset: 2px;
}

@media (prefers-reduced-motion: reduce) {
  .botao { transition: none; }
}

focus-visible instead of focus to not light up the ring on mouse click; outline-offset so the ring doesn’t die against the button fill; and the reduced motion rule, which also never comes alone. None of these three lines is difficult. All three depend on someone remembering to ask.

The layout works on the monitor that the model imagined

The classic, and what consumed the most chat rounds. The hero that came back:

.hero {
  height: 100vh;
  display: grid;
  place-items: center;
}

.servicos {
  display: grid;
  grid-template-columns: repeat(4, 1fr);
  gap: 24px;
  max-width: 1200px;
}

.depoimento {
  position: absolute;
  top: 420px;
  right: 80px;
  width: 380px;
}

Three traps in twelve lines, all of the same type: each value was chosen for a specific viewport that no one declared.

height: 100vh on mobile counts the browser’s address bar, so the content is cut off or jumps when the bar retracts—and if the hero text is larger than expected, it spills out of the block instead of pushing the page. repeat(4, 1fr) without a media query becomes four columns of 70px on a phone. And the position: absolute with top: 420px is the worst of the three because it works perfectly at the width the model thought of and falls apart at any other width, never giving a sign that something is wrong in the code.

The version that solves all three isn’t more complicated, it’s just written without assuming the screen:

.hero {
  min-height: 100svh;      /* svh respeita a barra do navegador móvel */
  display: grid;
  place-items: center;
}

.servicos {
  display: grid;
  grid-template-columns: repeat(auto-fit, minmax(15rem, 1fr));
  gap: 1.5rem;
}

min-height instead of height lets the content grow. auto-fit with minmax recedes on its own and dispenses with breakpoints. And the testimonial becomes a flow item, not a piece floating at a coordinate.

None of this is obscure knowledge, and the model knows everything: explicitly ask for “responsive grid without media query” and it returns auto-fit right away. It just doesn’t choose on its own because repeat(4, 1fr) is what appears in most tutorial examples—and tutorials are written for a screen capture screen.

Asking “make it pretty” always returns the same site

Here’s the part that I find really useful about this post, and the one that changed my way of working.

I asked to “make it prettier and more modern” at several points in the project. I received the same site every time. Not similar: the exact same one. Purple-to-blue gradient on the hero, white card with generous border-radius and diffuse shadow, Inter as the font, rounded icon above each service block, a row of three identical cards, and a dark footer. If you’ve seen any product landing page from the last three years, you’ve seen exactly this page.

And this isn’t laziness on the model’s part. It’s the correct behavior for it.

A language model returns the most likely continuation. “Pretty” and “modern” are adjectives without a referent—they don’t point to any value, any color, any measurement. Faced with an unrestricted request, the only path is the average of the training material, and the average of the training material for “pretty and modern site” is literally that: the visual dialect that dominated Dribbble and templates from 2021 to 2023. Asking for aesthetics without restrictions is asking for the dominant trend. You’ll receive the dominant trend, with precision, always.

The way out isn’t to write a better adjective. There is no better adjective. The way out is to stop sending adjectives.

Compare the two requests. First, what I used to send:

Deixa a página mais bonita e moderna, com uma cara mais profissional.

Now what I started sending—same model, same session, same input HTML:

Estilize com estas restrições. Não introduza nenhum valor fora delas.

Cores (só estas):
  fundo       #ffffff
  superfície  #f4f6f4
  texto       #1c211d
  secundário  #5f6b62
  acento      #1f5d3a   -> links, foco e botão primário
  marcação    #b23c17   -> só alerta, no máximo 2% da tela
  régua       #d9e0da   -> 1px, único separador do sistema

Tipografia:
  títulos  Archivo 700, entrelinha 1.06, tracking -0.03em
  corpo    Source Serif 4 400, 19px, entrelinha 1.72, coluna de 68ch
  mono     JetBrains Mono 13px, apenas em código e números

Espaçamento: use só 4, 8, 16, 24, 40, 64 e 104px. Nenhum valor intermediário.
Raio: 0, 2px ou 4px. Nada acima.
Sombra: nenhuma. Profundidade é régua de 1px ou superfície um degrau abaixo.
Estados: todo :hover tem :focus-visible equivalente, anel sólido de 2px.
Contraste: corpo no mínimo 7:1, secundário no mínimo 4,5:1.
Larguras: nenhum container com px fixo. Testar em 320, 768 e 1440.

Se algo do meu HTML não couber nessas regras, me diga em vez de inventar valor.

The second returns a site that is mine. Not because the model became more creative—it became less, and that’s what solves it. With the response space closed, it stops choosing between a million average pages and starts doing the only thing it does very well: applying a set of rules to a set of elements, with superhuman consistency. It doesn’t forget a spacing in twenty files. I do.

The trade-off is clear: aesthetics aren’t delegable; aesthetics application is. Someone needs to decide the palette, scale, and rules—that’s design work, and it doesn’t come out of a one-line prompt. After deciding, chat is the best cheap executor that has ever existed for that.

And the last line of the block is worth half of it. “Tell me instead of inventing value” trades a silent decision—which is expensive because you only find out later—for a question, which is cheap.

Accessibility is the hole that doesn’t make noise

Combine what has appeared so far and notice the pattern: div instead of nav, 4.47:1 contrast, hover without focus, outline: none, height: 100vh. All are accessibility failures. And none of them manifest.

Syntax error breaks the build. Logic error breaks the test. Accessibility error does nothing: the page opens, renders, looks nice in the screenshot, the client approves, the site goes live. Feedback only comes through a path that most projects never take—someone navigating by keyboard, someone using a screen reader, someone with low vision on their phone in the sun, or an audit.

That’s what I did at the end, and it’s the step I would recommend to anyone who builds a site this way. A pass of Lighthouse and one of axe DevTools, both free and within the browser, found in a few minutes a list of problems that I hadn’t seen in weeks looking at the page. After that, three manual tests that cost two minutes each and catch what automated tools don’t:

  • navigate the entire page only with Tab and verify if you can see where the focus is all the time;
  • zoom to 200% and check if any text disappears or horizontal scrolling appears;
  • narrow the window to 320px wide and see what breaks.

None of these three requires specialized knowledge. They just need someone to remember to do—and it’s exactly this “remember” that chat doesn’t do for you because it responds to what was asked and no one asked.

Where this is still worth it

I ended up with a working site, and I don’t regret the method. I just know better where it applies.

It’s very useful for prototyping. When the goal is to see the idea on screen to decide if it survives, the average site is great—it’s average precisely because it’s familiar, and familiar is what you want in a demonstration. Nothing there goes into production, so accessibility debt doesn’t even arise.

It’s useful for disposable landing pages. Page of an event that dies in three weeks, registration form for a class. Zero maintenance cost because there is no maintenance.

And it’s useful, with a serious caveat, for those who aren’t front-end and need something that works. A back-end that needs an internal panel, someone building their own business site. Here the chat delivers, in an afternoon, something that this person wouldn’t build in a week. The caveat is that they also don’t have a way to see any of the problems in this post—and that’s why automated auditing stops being good practice and becomes mandatory. Two tools, ten minutes, for free.

Where I wouldn’t use it like this: anything with its own brand to defend, anything another person will maintain, anything where accessibility is a contractual requirement. Not because the model can’t handle it—because in these cases the hard work is deciding the restrictions, and that part remains mine.

What’s left at the end is a division of labor rule that I didn’t have before. The model is excellent where there is a correct answer and terrible where there is only an adequate answer—and the difference between the two isn’t technical difficulty, it’s the presence of a context that only I have. HTML has a template. CSS has a project. As long as I send adjectives, I receive the average of the internet; when I send hexadecimal, scale, and rule, I get my site—assembled by a machine that doesn’t make the twentieth application of the same rule wrong, which is precisely where I go wrong.