How Much of the Industry Can We Actually See?

Coverage of the starlikeness index is narrow and measurable: 241,792 face vectors, but only 2,333 identity records, of which 1,707 carry a real profile. Roughly 1% of indexed faces are linked to a name, and the index measures aggregator exposure rather than industry output.

Last updated Mon Aug 03 2026 00:00:00 GMT+0000 (Coordinated Universal Time)

Less than you would guess from the headline number. Our index holds 241,792 face vectors but only 2,333 identity records, and only 1,707 of those carry a resolvable profile with biographical fields. About 1% of the faces we have indexed are attached to a name (our index, 2026-08 snapshot).

That ratio is the honest answer to "how complete is your data". Everything else in this article is an attempt to say precisely which 1%, and why the other 99% is not a failure of the system but a property of the material.

What do the headline numbers actually count?

Each number counts something narrower than it sounds, so here is each one with its real definition.

Figure Value What it counts What it does not count
Face vectors 241,792 Faces detected in indexed images Distinct people; one person appears many times
Identity records 2,333 Clusters treated as one person with a label Everyone in the index; most faces are unlabelled
Records with a full profile 1,707 Identities with a canonical slug and metadata The 626 that hold only a raw name string
Sites indexed 106 Hosts we have crawled imagery from Sites carrying Japanese content IDs — around 13
Keyword landing pages 300 Query pages built over the index Coverage of the industry's tag vocabulary
Locales 13 Interface and name-translation languages Sources; most sources are English-language

Source: our index, 2026-08 snapshot.

The first row is where most face-search marketing lives, and it is the least meaningful one. 241,792 is a measure of how many images we processed, not of how many people we can identify.

Why is only 1% of the index named?

Because faces are found mechanically and names are supplied editorially, and the editorial supply is thin on the sources we index.

A detector finds every face in every image we process. A name only attaches when a source page states it in a parseable, linkable form. On international aggregator sites, which make up the bulk of our 106 hosts, credits are often stripped entirely during republication.

The consequence is visible inside our own similar-performer data: of 13,420 neighbour references, 7,004 (52%) point at identities with no name. More than half of all "who resembles this person" pointers lead somewhere the public record left blank.

How complete is the data on the people we have named?

Between 23% and 89% depending on the field, and we publish the low figures alongside the high ones.

Field Records with a value Coverage of n=2,333
Similar-performer list 2,070 89%
Cup size 1,375 59%
Debut date 1,371 59%
Height 1,340 57%
Measurements 1,316 56%
Birth date 1,301 56%
Birthplace 716 31%
Blood type 547 23%

Source: our index, 2026-08 snapshot, n=2,333 identity records.

Two things follow. First, any statistic we could publish about, say, birthplace distribution rests on 716 people — and those 716 are not a random sample, they are the ones whose profiles were best documented, which usually means the most prominent. Second, a field like blood type at 23% is too thin to support a population claim at all, and we would rather say so than dress 547 records up as an industry portrait.

The one strong field, similar-performer lists at 89%, is strong precisely because we compute it ourselves from vectors rather than inheriting it from a source.

What does the index measure, if not the industry?

It measures exposure on aggregator sites, and the distinction is not academic.

2,246 of our 2,333 identity records take their representative image from a single site. Across the entire set, representative images come from only four distinct hosts. Of the 106 sites we index, only around 13 carry Japanese content IDs; the remainder are largely international tube and aggregator platforms.

So when a performer is well represented in our data, the finding is: this person's imagery circulates widely on the aggregators we index. That is a real and useful signal. It is not the same as being prolific, popular in Japan, or commercially significant, and we do not present it as any of those.

This also bounds what we will never claim. We will not say a studio released N titles — only that we have indexed N. The index has no release-date field, no studio field, and no title records at all, so year-on-year production trends, studio market share, and format-adoption timelines are simply not computable from it, regardless of how confident a chart made from it would look.

Data and method

Source. The starlikeness face index: face vectors and identity records derived from cover images and video stills across 106 sites.

Sample. 241,792 face vectors; 2,333 identity records (1,707 with a full profile, 626 name-string only); 13,420 similar-performer references.

Method. Detection with YuNet; alignment from five facial landmarks; encoding with ArcFace into 512-dimension vectors; retrieval by cosine similarity with a 0.40 same-person threshold. Coverage percentages are the count of records holding a non-empty value for a field, divided by the full population of 2,333.

Snapshot date. 2026-08. Counts recomputed from the live data file on that date; the visible update date on this page reflects the data, not a rebuild.

Known bias.

  • Source concentration. 2,246 of 2,333 representative images come from one site; four hosts account for all of them. Anything absent from those sources is absent from our measurements.
  • Language and region skew. Most indexed sites are international aggregators, so the imagery that circulates internationally is over-represented relative to material distributed only in Japan.
  • Documentation skew. The named 1,707 are, by construction, the best-documented performers. Coverage statistics computed over them describe the well-documented, not the typical.
  • Historical prior. Some field coverage figures previously published on this site were inflated by placeholder strings being counted as values. Those were removed before this snapshot; the figures above are the corrected ones, and they are lower than what we reported before.

Citation.

starlikeness (2026). "How Much of the Industry Can We Actually See?"
Face index, 2026-08 snapshot (n=241,792 faces; 2,333 identity records;
106 sites). https://starlikeness.com/en/posts/how-complete-is-our-index

Related questions


The practical test of coverage is a search. Upload a frame and read the scores: a list of low-scoring results is the index telling you this person is outside the 1%. Detection runs in your browser and the original image never leaves your device.

Frequently asked

Does a large face count mean broad industry coverage?
No. Face count measures how many faces were detected in the images we indexed, and a single production contributes many faces. Coverage of people is measured by identity records, which in our case is 2,333 — two orders of magnitude smaller than the face count.
Can you calculate how many titles a studio released?
No. The index stores faces, not titles, and carries no studio, release date or title-count field. Any number derived from it would describe what aggregator sites republished, not what a studio produced.
If a performer is missing, does the search say so?
Not directly. Dense vector search always returns its nearest results, so absence appears as a list of low scores rather than an empty result. The score is the signal; there is no "not found".