Skip to content

Glossary HTML

What Is HTML?

Definition

HTML (HyperText Markup Language) is the language used to set out the structure and meaning of a web page's content: what is a heading, what is a paragraph, what is a link or an image.

On this page 5
  1. What HTML means
  2. How it works
  3. Why it matters
  4. Best practices
  5. Common mistakes
In brief

HTML gives a web page its structure and meaning: what it is, how browsers and search engines read it and what AI crawlers actually see of it.

What HTML means

MDN describes HTML as the most basic building block of the web: it defines the meaning and structure of content. Appearance is handled by CSS, and interactive behaviour by JavaScript. HTML is a markup language: it labels each part of the text with tags that say what it is, without calculating or deciding anything.

Each tag sits between angle brackets. The h1 tag opens the main heading and its version with a slash closes it; the p tag marks a paragraph; the a tag, with the href attribute, creates a link. Tags can carry attributes with extra information, such as a link's destination, an image's alternative text or the page's language.

The standard in force is the HTML Living Standard maintained by WHATWG, a document that is updated continuously: the latest published version is from 7 October 2026. There are no numbered versions any more; changes go into the living standard as browsers adopt them.

How it works

When someone opens a page, the browser downloads the HTML from the server, reads it from top to bottom and builds an in-memory structure from it, the DOM. Along the way it finds references to other files, such as stylesheets, scripts and images, and downloads those too. With all of it, it draws the page on screen.

A page can reach the browser in two ways. With server-side rendering, the HTML already contains the full content. With client-side rendering, the server sends an almost empty HTML file and JavaScript fills in the content afterwards, in the browser. For the person looking at the screen, the result can be identical. For a crawler that does not run JavaScript, it is not: in the second case, the page is empty.

Search engines read HTML for a different purpose than browsers. Their aim is to understand what the page says and how it relates to others: the title and meta description, headings, text, links, the canonical tag, language alternatives (hreflang) and structured data.

Check it in the page source, not in the browser window.

Why it matters

For SEO, the HTML is the version of the page that really counts. A heading that only exists as an image, text that loads on a click or a link that is a button without an href stays invisible or unclear to a crawler, even if a person understands it without trouble.

In AI search, the difference is even bigger. In December 2024, Vercel published an analysis of crawler traffic on its network and concluded that none of the major AI crawlers, including those from OpenAI, Anthropic and Perplexity, ran JavaScript; only Gemini, which uses Googlebot's infrastructure, and Applebot did. Whatever is not in the initial HTML does not exist for those crawlers.

We measured this in this very glossary, on the entry about CMS, on 8 October 2026. The server delivers 186,694 bytes of HTML, and the readable text takes up 12,769 of them, less than 7%. The rest is structure, component styles and scripts. What matters is that those 2,049 words, including the five frequently asked questions, are already in the HTML that arrives without running anything, together with a single h1, fourteen h2s, the lang attribute set to "es" and four hreflang references (es, en, de and x-default). That is what a crawler that does not run JavaScript reads.

Best practices

  • Check with "view source" or a request without JavaScript that the main content is in the HTML the server delivers.
  • Use one h1 per page and a heading hierarchy that reflects the real structure of the text.
  • Write links as a tags with an href and a real URL, not as buttons or JavaScript events.
  • Declare the language with the lang attribute and, if there are several versions, link them with hreflang.
  • Give images that carry information a useful alternative text.
  • Validate the markup from time to time: a missing closing tag can cause a whole block to be read wrongly.

Common mistakes

  • Loading the main text only with JavaScript and assuming every crawler will see it.
  • Using heading tags to make text bigger instead of to mark its hierarchy.
  • Hiding important content behind tabs or accordions that only fill in on a click.
  • Repeating the same title and meta description across many pages.
  • Confusing what the browser inspector shows, already changed by JavaScript, with the HTML the server sent.
Manuel Riveiro Rodriguez CEO & Digital Strategist

A technical audit covers this and everything else in one pass.

Request an audit

Frequently asked

Is HTML a programming language?

No. HTML is a markup language: it states what each part of the content is, such as a heading, a paragraph or a link, but it does not calculate anything or make decisions. JavaScript does that, and CSS handles the visual presentation. The three work together on almost every web page.

What is the difference between HTML and CSS?

HTML says what each thing is and in what order it appears. CSS says how it looks: colours, fonts, sizes and position. You can change a page's appearance completely by editing only the CSS, without touching its HTML, and the content search engines read would stay the same.

How does HTML affect SEO?

It is what search engines read to understand a page: title, headings, text, links, canonical, language and structured data. If the important content is not in the HTML or is marked up wrongly, the search engine understands it less well or does not see it at all. Clean HTML does not guarantee good rankings, but faulty HTML can prevent them.

Do AI crawlers read a page's JavaScript?

According to the analysis Vercel published in December 2024, the major AI crawlers, such as those from OpenAI, Anthropic and Perplexity, did not run JavaScript; Gemini did, because it uses Googlebot's infrastructure. That is why the content should already be in the HTML the server delivers.

What is HTML5?

It is the name under which the HTML version became known that introduced, among other things, semantic tags such as header, nav, article and footer, and native support for audio and video. There are no numbered versions today: WHATWG maintains the HTML Living Standard, which is updated continuously.

Sources

  1. MDN Web Docs, "HTML: HyperText Markup Language". Source for the description of HTML as the most basic building block of the web, defining the meaning and structure of content alongside CSS and JavaScript.
  2. WHATWG, "HTML Living Standard", last updated 7 October 2026. Source for the standard in force and its continuous updating.
  3. Vercel, "The rise of the AI crawler", 17 December 2024. Source for the finding that the major AI crawlers analysed did not run JavaScript, with Gemini and Applebot as exceptions.