Tracking down Life Overlooked
Data Deficient status on its own predicts very little. Of the species here that IUCN has since reassessed, the group as a whole came out threatened at roughly the ordinary all-species rate. The narrowly restricted subset is different, and that distinction, not the Data Deficient label, is what this list is for.
This is a ranked subset, not a census. IUCN v2026‑1 holds 8,659 Data Deficient species in vertebrate classes; this list covers 3,031 of them, which is 35%. The method needs a published sentence describing each species' range, and the corpus behind it is one reference work. Of the 2,964 Data Deficient vertebrates that corpus names, every one is in the list; the other 5,628 are never named in it at all. So the right reading is "of the Data Deficient vertebrates with a usable published range statement, these are the most narrowly restricted", not "these are the ones that matter most".
How long is a blank allowed to stand? The Red List export says only that a species is Data Deficient now. Its assessment history, pulled from IUCN's API, says since when, and the answer reframes the list. The median species here has been Data Deficient for 16 years; 475 have held the label for twenty years or more. And 1,210 of them, 40%, have been assessed more than once, meaning IUCN returned, looked again, and wrote Data Deficient a second or third time. Those are not species awaiting a first look.
For scale, the median time IUCN has actually taken to resolve a Data Deficient listing is 16 years, so 1,011 species in this list are already past it.
The species outside were screened out, not overlooked. Richardson takes an explicit position on the Data Deficient category: most of it, in his judgment, is junior synonyms or valid species not in immediate danger, while genuinely threatened species get lost there for want of follow-up research.
He admits a Data Deficient species when its record is narrow enough to imply risk, and that screen is measurable, since he kept 28% of IUCN's Data Deficient ray-finned fishes against 72% of the birds, dropping the deep-sea groups almost entirely because naturally rare is not the same as threatened. So the omitted species are neither unexamined nor shown to be safe: one author's judgment, made without the research he says the category lacks, is not an assessment. It also means this corpus is pre-selected for the trait the list ranks on, which is why 92% of records state some restriction. The list tests an expert's screen against IUCN's later verdicts rather than discovering the pattern from scratch.
Nobody with a specialist's authority. The ranking is produced by a text-classification rule that reads the source's own range statements and sorts them by severity. No taxonomist or conservation biologist has reviewed the individual species. What has been tested is the rule's average behavior against IUCN's independent reassessments, a statement about the group, not about any one row.
These are not Red List assessments and must not be cited as such. They are a prioritized research agenda: an argument about where to look.
Target 4 of the Kunming-Montreal Global Biodiversity Framework commits governments to halting the extinction of known threatened species, a phrase that excludes both categories in this list by definition. Funding follows the same line: under 3.5% of Mohamed bin Zayed Species Conservation Fund awards have gone to Data Deficient taxa, and under 1% at the People's Trust for Endangered Species. IUCN's own guidance says Data Deficient species should get the same attention as threatened ones, and it is not followed.
So the consequence of having no category is not merely that a species is poorly known. It is excluded from the target that mobilizes national action and from the funds that pay for fieldwork, which is what makes deciding which of these species to look at first worth doing carefully.
| Evidence of risk | Matthew Richardson, Threatened and Recently Extinct Vertebrates of the World: A Biogeographic Approach, Cambridge University Press, 2023. Vertebrates only, the scope follows the book's. |
|---|---|
| Extinction risk | IUCN Red List v2026‑1, 178,011 taxa, via GBIF's Darwin Core mirror (accessed 2026‑07‑28). |
| Taxonomy | GBIF Backbone, for synonym resolution and for independent verification of assessment status. |
| Map points | GBIF georeferenced occurrence records, queried by strict taxon key. Basemap is Natural Earth 110m land (public domain). |
Every species is assigned one of five restriction tiers. The tier records one thing only: how narrow the species' published record is, read from the range statement in Richardson's own text. Tier 5 means the species is known from a single specimen or nothing beyond its type locality; tier 1 means the source makes no restriction claim at all.
What a tier is not: it is not an IUCN category, not a threat level, not a measure of how endangered the species is, and not a count of individuals. Two species in the same tier may face completely different pressures. The tier is a statement about the evidence, not about the animal.
Restriction is the axis because IUCN Criterion B can qualify a species as threatened on a small range: area of occupancy below 2,000 km² for Vulnerable, 500 for Endangered, 10 for Critically Endangered. The area figure alone never qualifies anything. Criterion B also requires at least two of three further conditions, which are severe fragmentation or very few locations, a continuing decline, and extreme fluctuation. A species known only from its type locality falls inside the area figure for Critically Endangered and is silent on the other two conditions, which is exactly why a tier here is a prompt to assess rather than a category. That is what a narrow record is worth ranking on. Tiers 3, 4 and 5 together are the priority stratum, and they are the only part of the ranking that has been tested. The boundary sits there because that is where the tested risk actually jumps, tier 2 came out threatened at 20%, tier 3 at 67%, and not because the top of the scale is the most severe: it is not, see Methods.
| Tier | Basis | Species |
|---|
No area of occupancy has been computed. A priority flag argues that Criterion B should be evaluated; it is not a finding that it is met.
The strongest objection to ranking on range comes from the Red List itself. Fifteen of its architects, including Red List Committee and Standards and Petitions members, wrote that "the Red List classifies extinction risk rather than rarity" and that rarity "does not consistently lead to high extinction risk", describing extent of occurrence as an index of insurance against threats rather than a depiction of a species' range (Collen et al. 2016).
The nearest empirical test cuts both ways, and it comes from the same research group. Geyle et al. (2020), whose senior author is Chapple, put the 60 Australian squamates already in the highest IUCN categories through structured expert elicitation. Every one of the 20 most imperilled turned out to be restricted in range, which is the pattern this list is built on. Their finding, though, is about timing: six of the 60 carry a greater than 50% chance of extinction within 20 years, and up to 11 could be lost in that window without better management. Restriction and imminence are different measurements. A narrow record marks a species worth assessing, and says little on its own about how much time is left.
There is also a direct precedent worth knowing. In 2016 a study argued that many forest birds warranted uplisting once mapped range was replaced by estimated suitable habitat, and IUCN's Red List Committee, Standards and Petitions Sub-committee and BirdLife International replied formally that the approach conflated two different thresholds and was "not conservative, but drastic". That ruling is why nothing here assigns a category. This list says a species' record looks narrow enough that somebody should evaluate it. It does not say the criterion is met, and it substitutes no range measure into any threshold.
The priority rate lands close to the 56% that Borgelt et al. (2022, Communications Biology) predicted for Data Deficient species generally. One reading is that the figure belongs to the narrowly restricted subset, and that an unranked list dilutes a real signal with vague-range species. That reading is worth stating and worth doubting in the same breath: Borgelt's threatened predictions carry a precision of only 60% to 67%, and the best-audited reptile study found no elevation among Data Deficient species at all. The section below sets out both sides.
Two further caveats are structural, and they bound what this test can mean. The first is circularity. The ranking is built on how narrow a species' published record is, and IUCN assesses many of these species under Criterion B, which is itself a range criterion. So part of what the test measures is agreement with the criterion an assessor was already likely to reach for, rather than an independent prediction of extinction risk. That does not void the practical claim, which is about what an assessment yields, but the result should not be read as a biological finding about narrow-range species. The second is range underestimation. Cazalis et al. (2023) found that Data Deficient status is predicted mainly by how few records a species has rather than by its ecology, and warned that a small known range for such a species may partly reflect that nobody has looked. On this list the two cannot be cleanly separated, because the range and the evidence of effort are read from the same published sentence.
The premise behind this list, that Data Deficient status conceals threatened species, is contested rather than settled. A review of work published between 2021 and 2026 is summarized here because anyone deciding whether to trust this ranking should see the evidence against it beside the evidence for it.
Supporting it. Borgelt et al. (2022) predicted 4,336 of 7,699 Data Deficient species, 56%, to be threatened against 28% of data-sufficient species, though the paper reports precision of only 60% to 67% for its threatened predictions. Isa et al. (2024) supply the strongest real outcome rather than a prediction: of 656 amphibians once listed Data Deficient and later reassessed, 341, or 52%, came back threatened, against 40% of amphibians never listed Data Deficient. That is an elevation of roughly 1.3 times, not a doubling. Chapple et al. (2026) report that 16 of 42 Data Deficient Australian squamates, 38%, carry evidence of being threatened against 7% of data-sufficient species, and that only 6 of those 42 attracted any published ecological work in the eight years after assessment.
Against it. The sharpest counter-evidence is Bland and Böhm (2016), who modeled every Data Deficient reptile in the Sampled Red List Index and found 19.2% predicted threatened, statistically indistinguishable from 19.1% for data-sufficient reptiles. When 60 of those species were later genuinely reassessed, 8 became threatened, 13%, which is below the baseline rather than above it. That position is also IUCN's own: the Global Reptile Assessment derives its headline of 21.1% threatened by assuming Data Deficient reptiles are threatened at the same rate as assessed ones. Cazalis et al. screened 6,887 Data Deficient species for habitat loss and flagged 112, or 1.6%, as plausibly qualifying on that criterion. And Meiri et al. (2018) examined 927 lizards known only from a type locality, the same population this list calls tier 5: 213 are known from one specimen and 736, 79%, were never recorded again. Their title is "Extinct, obscure or imaginary."
What the models get wrong is specific. Every validated study found the same asymmetry. Di Marco (2022) reports that automated predictions were right about 92% of species that later came back Least Concern and under 30% of those that came back threatened. Models are good at confirming safety and poor at catching danger, which is the opposite of what a triage tool needs, and it is the error this list should be expected to make too.
The size of the disagreement is the finding. Reviewing the earlier window, 2016 to 2021, makes the spread plain. The same quantity, the share of Data Deficient species that are really threatened, has been put at anywhere from 8% to 69% depending on the method used (Jarić et al. 2016). Sharks and rays show the instability inside one taxon: Data Deficient species came out above the assessed baseline in the northeast Atlantic, level with it in the Mediterranean, and below it globally (Walls and Dulvy 2020), whose models did not find range size useful at all. A number that moves that much across methods and regions is telling you about the methods.
The best answer to the circularity objection. A meta-analysis of 173 studies across all kingdoms found range size to be the trait with the most significant tests, stated the circularity risk plainly, and then tested it: even after excluding species listed as threatened because of a small range, range was still strongly associated with extinction risk (Chichorro, Juslén and Cardoso 2019). That is the strongest rebuttal available, and it carries one limit worth stating: it concerns data-sufficient species whose ranges are observed, not Data Deficient species known from a handful of records.
What almost nobody has measured. No published study reports the comparison this list actually makes, narrow-range Data Deficient species against broader-range Data Deficient species, checked against later confirmed assessments, at any scale beyond the 140 tested here. Reviewing both five-year windows found no analogue in either. That cuts two ways: the test behind this ranking is of a kind nobody else has run, and it is therefore also unreplicated. There is likewise no published figure for how long a species stays Data Deficient, and no Data Deficient risk estimate for plants at all. And no published model in either window includes survey effort as a covariate, though at least 80% of biodiversity records for every taxon studied lie within 2.5 km of a road (Hughes et al. 2021). Since that is the strongest objection to this list and nobody had run it, we ran it here. The result is in the next section.
So the honest position is that this ranking is a hypothesis with two small tests behind it, pointing the same way as some of the literature and against the best-audited reptile study in it.
The odds ratio, with its interval. The 4.64 above comes from 15 of 27 against 24 of 113. Its 95% confidence interval, by conditional maximum likelihood, runs from 1.74 to 12.33, with a Fisher exact p of 0.0007. The direction is solid, because the interval excludes 1. The magnitude is not, because it spans a factor of seven. The defensible claim is the floor, not the point estimate.
The effort objection is real in this data, and the signal survives it. Narrowly recorded species here do have fewer occurrence records, a median of 7 against 19 for the broader-range group, which is exactly the entanglement Cazalis et al. describe. Adding the base-ten logarithm of each species' GBIF record count to the model moves the odds ratio from 4.64 to 4.14 (95% CI 1.68 to 10.21, p = 0.002). That is a shrinkage of 11%, well inside the noise, and record count on its own carries no significant signal (p = 0.22).
Stratifying is more transparent than modelling at this sample size. Splitting the labelled set into thirds by record count gives odds ratios of 3.05, 4.00 and 6.20 from the least to the most surveyed third, pooling to 3.88 (Mantel-Haenszel, 95% CI 1.55 to 9.70). The artifact explanation predicts the opposite shape: a large ratio among the poorly recorded species and nothing left among the well recorded ones. That is not what this sample does. On this evidence restriction is not simply a proxy for how much anybody has looked.
One suggested refinement was tested and did not hold. Meiri et al. propose that a species described long ago and still known from one locality is more likely to be genuinely restricted than a recently described one. Split at the median description year of 1966, the odds ratio is 2.65 among the older names and 8.64 among the newer ones, which runs opposite to the proposal, and the interaction term is not significant (p = 0.25). With eight species in the smallest cell the honest reading is that this cannot be measured here, rather than that it is absent.
The tier is a rule over Richardson's wording, so it reads a phrase and cannot see a depth. "Known only from a single specimen" is tier 5 whether the specimen came off one hillside or out of one net tow in the bathypelagic, and those two sentences do not mean the same thing. Ray-finned fishes are half this list, so the objection is not a technicality.
So the ranking is reported as a claim about land and fresh water. Oceanic species remain listed, because a list that hides its weakest stratum is worse than one that labels it, but the validated signal does not extend to them and no priority claim here rests on them. Use the realm filter on the list to see either group on its own.
A locality is recorded two ways, and the Localities page distinguishes them.
From a locality account. Richardson organizes the book geographically, so most species sit inside a paragraph introduced by a place. That place, and its opening sentence, are attached to every species in the paragraph. This is the stronger provenance and it covers 518 species.
From the species entry itself. Where a species is not inside a locality account, the place is very often named in its own sentence anyway: "collected from the upper Pungwe River", "collected off Mauritius". Pattern extraction over those sentences recovers a locality for a further 1,211 species, taking coverage from 19% to 60% of the list and the site list from 42 to 102. Clusters that were entirely invisible before appear this way, the Gulf of California holds 19 species, none of which had a locality until now.
Named physical features (Salween River, Gulf of Panama, Mount Kinabalu) are preferred over a bare capitalized word after a preposition, and a capture is rejected if it collides with the species' own common name or genus. The place is only ever taken from the text, never guessed, if the sentence names no place, the species keeps no locality, which is why 1,333 still have none.
These names are then looked up in a gazetteer to place them on the map, which is a separate step with its own failure modes and its own verification, see Geocoding the locality names below. Nothing in the list or on the Localities page depends on that lookup: those show the name the source gave, and the coordinates exist only for the map.
The list stores localities as text taken from Richardson's prose, "The Mano River", "Arrowsmith Bank", "Espiritu Santo". Those examples are the specific end of a range: the field carries whatever resolution the book gives, so it also holds section headings and habitat descriptions. Of the 1,729 species with a locality string, fewer than half are specific enough to place at all, and the largest single value is "The Tropical Atlantic Region", covering 63 species. That is faithful to the source and coarser than the three examples above suggest.
Those strings were not run through a gazetteer to make the map: place-name geocoding is ambiguous exactly where this corpus is ambiguous, and a map that confidently places a species in the wrong ocean is worse than no map. The points are real museum specimens and field sightings that already carry map coordinates (“georeferenced” records), taken from GBIF and queried per species.
One trap governed that query, and it is the same species of error that wrecked the first never-assessed list. GBIF's free-text name search must not be used, because many of these names are synonyms and it silently returns records of the accepted species instead:
| Anampses viridis, free-text search | 3,071 records | centroid inland in New South Wales |
|---|---|---|
| Anampses viridis, strict taxon key | 3 records | all at Réunion |
The book describes that species as known from a single specimen; the 3,071 belong to Anampses caeruleopunctatus, its widespread senior synonym. Every name is therefore resolved to a GBIF usage key by strict match first, and records are pulled by taxon key. Records GBIF itself flags as having bad coordinates are excluded, and points are capped at 30 per species so well-collected species cannot dominate the surface.
What the map is not. Not range maps, these are point records. Not a picture of where risk is concentrated, it maps collecting effort at least as much as biology, because georeferenced records cluster where institutions and expeditions have worked. And not the same basis as the Localities page, which comes from the book rather than from GBIF, so the two will not correspond exactly. Some bad coordinates survive the filtering, which is why the page shows density rather than authoritative dots.
The list's place names are also plotted, as rings, because refusing to look up "Salween River" while happily printing it was inconsistent. But a gazetteer answers any string, and two checks are needed, not one.
Type agreement catches the obvious failures: asked for "The Cordillera Central", Nominatim returns a university in Baguio. A name ending "River" must come back as a waterway, "Island" as an island; buildings, roads and campuses are rejected outright. That removed … of … candidates.
But type agreement cannot see a same-type homonym, and those are common. "Congo River" resolved to a real river of that name in Sierra Leone, 4,000 km from the Congo, passing cleanly as river/river. "Dunk Island" resolved near Sydney rather than Queensland. So each geocode is then measured against the GBIF records of the species that sit at that locality, an independent witness already held:
A geocode that fails that test is not simply dropped. Disagreement says the coordinate and the records are at odds, not which one is wrong, so the failures were worked back through. Asking the gazetteer only about the country the species lives in recovered 13. Asking it only about the box its own records sit in recovered 13 more, Congo River included: the Sierra Leone river cannot come back at all when the question is restricted to the Congo basin. Reading the remainder by hand recovered 8 and showed 16 had been right all along, refused only because the feature or the species is bigger than 500 km.
| Agree within 500 km, plotted | — |
|---|---|
| Disagree by more, discarded | — |
| No species with records to check against, off by default | — |
Around one in six checkable geocodes was wrong by more than 500 km, and none of those errors was detectable from the name or the feature type alone. That rate is the reason the unverified … are not plotted by default: there is no reason to think they are cleaner, and they cannot be tested. The validated rings add … species that have no georeferenced record of their own, modest, but every one is corroborated.
A locality naming two places is kept but never plotted. … strings in the list are compounds , "Athi and Tana River", "Caroline and Marshall Islands". There is no single coordinate for them, and every way of reducing one to a single place either drops the noun that identified it ("Caroline") or manufactures somewhere that does not exist ("Oman and Masirah Island" → "Oman Island"). They are left whole, excluded from the gazetteer, and the evidence sentence shows both places.
The most telling number is the one that cannot be plotted. Of the 3,062 species here, 805 have no georeferenced record anywhere in GBIF, and 726 cannot be placed even by a cross-checked locality. They are the least visible members of an already invisible group, and they are absent from the map.
Two fields that look like risk signals are recorded on each species and deliberately kept out of the ranking: whether it has gone unrecorded recently, and whether a threat is named at its locality. Combined with restriction they produce a non-monotonic score, and naming a site-level threat is inversely associated with being threatened. The reason is legible: broad, well-documented regions attract threat prose but hold wide-ranging species, while a genuine type-locality endemic gets one terse line. Both fields stay on the record as description, and neither is ranked on.
Proving that IUCN has never assessed a species is much harder than it looks, because the two sources disagree about naming constantly. Deciding it by failing to find a name produced a false-positive rate near 90%. Four distinct failure modes had to be closed:
| Failure | Example |
|---|---|
| Genus reassigned | Amblyopsis rosae is IUCN's Troglichthys rosae (Near Threatened) |
| Genus and family reassigned | Aphanius iberus is IUCN's Apricaphanius iberus |
| Homonym within an order | Sinosuthora przewalskii, an Asian parrotbill, matching Grallaria przewalskii, a South American antpitta |
| Rank change | Cyanoramphus cookii, in fact Least Concern |
Species known only from pre-1500 remains were also removed. The Red List covers extinctions from 1500 AD onward, so those fall outside its remit by policy rather than by neglect, counting them as coverage gaps would be a category error.