natzir.com

Schema.org · public usage dataset · may 2026

What the web
actually marks up.

In May 2026, Google and Schema.org published how often each Schema.org term is used across the public web. The data counts unique domains, grouped into popularity buckets. The vocabulary holds 5,545 terms, yet most adoption sits on a few dozen of them.

Terms measured

5,545

Types + properties

Broadly adopted

43

Used by more than 10 million domains

The long tail

77%

Used by fewer than 1,000 domains

01

Ranked by adoption

Six adoption tiers

Every term falls into a domain-count bucket. Sorted from most adopted to least, the bottom tier (fewer than 1,000 domains) holds more terms than all five other tiers combined.

Bar length ∝ √(number of terms), rescaled per view. Counts are unique domains, not pages or URLs.

02

The head of the curve

The most-used terms

The most-adopted terms are structural and commercial: breadcrumbs, websites, organizations, products, offers, ratings, articles. These are the terms found on the widest range of domains across the web.

Types

top 47

Used by 10M+ domains · 12

Used by 1M–10M domains · 35

Properties

top 96

Used by 10M+ domains · 31

Used by 1M–10M domains · 65

03

The reality check

The long tail

This curve doesn't measure what's useful. It measures what Google rewards.

We mark up what earns rich results and ignore the rest. Of Schema.org's 4,587 properties, 3,779 sit below a thousand domains; most of the vocabulary describes things no rich result depends on. Even the well-known rich-result types rank in the middle of the table, well below the structural and commerce terms:

04

All 5,545 terms

Explore the full dataset

Search any term, filter by kind or adoption tier, sort by name or popularity. Each term links to its Schema.org definition.

Schema.org terms with their kind and domain-adoption tier. Sortable by column.
05

Read before you cite

How to read this data

01

Domains, not pages

A term is counted once per domain. Product used on 100,000 pages of one site still counts as a single domain, so the number reflects how many sites use a term, not how often it appears.

02

Formats are merged

JSON-LD, Microdata and RDFa are counted together. The data does not say which syntax a site used to write a term.

03

Google's index only

These are counts from Google's public crawl. Sites blocked in robots.txt are not included, and the numbers reflect what that crawler sees.

04

Small does not mean unused

A low count is not always low value. Medical and government terms serve a narrow part of the web, so they sit in the lower tiers even when they are standard within their field.

Tiers 10M+ 1M–10M 100K–1M 10K–100K 1K–10K < 1K

Each tier is a range of unique domains, not an exact count. Grouping this way smooths daily crawl noise and protects site privacy.

06

Where it goes next

A living, open dataset

The dataset is refreshed every month, and its format is open for anyone to use or extend.

01

Refreshed monthly

A new snapshot is published to GitHub each month, after manual validation. Adoption changes slowly, so one snapshot is enough to see roughly where a term stands.

02

Open to other crawlers

The format is documented and not tied to one provider. Other search engines and web archives can publish their own counts in the same structure, which would extend the picture beyond a single crawler.

03

Built to act on

Use it to see which markup has real reach, and to check SEO tools and plugins against published adoption numbers instead of guesswork.