AI Can Describe the Manzanita. Somebody Still Has to Go Outside.
The chatbot has read the field guide. The hillside remains unimpressed.
Ask an AI about a manzanita shrub and it will hand you a polished little botanical sonnet: waxy leaves, twisting branches, red bark, drought tolerance. It has digested field guides, herbarium records, ecological papers, captions from photographs, and the accumulated prose of people who have spent serious time squinting at plants.
What it has not done is stand on a fire-scarred hillside, notice a weird pattern on the leaves, learn from a local ecologist that it appeared after the burn, take a usable sample, and decide whether this is disease, adaptation, or just the plant having a bad year.
That gap is not a sentimental argument for humans. Humans are frequently wrong, biased, distracted, and equipped with a clipboard they forgot in the car. It is an argument for something more annoying and more necessary:
Reality does not enter a database by magic.

The upload failed because the hillside declined to become a PDF.
AI is extremely good at rearranging the evidence already collected. It can draft, translate, summarize, code, generate images, suggest experiments, scan medical images, model weather, and identify promising chemical structures at a speed that makes ordinary office work look like a hand-cranked butter churn.
But cheaper output is not cheaper knowledge.
The archive is not the world, despite its very confident tone
Every model is built from an archive: text, images, sensor records, labels, experiments, code, and the residue of whatever institutions managed to notice, measure, digitize, and keep.
That is a spectacularly useful pile of stuff. It is also a biased pile of stuff.
It reflects:
-
What somebody had money and permission to study
-
What was safe enough to collect
-
What was considered worth publishing
-
What got recorded in a major language
-
What survived the filing cabinet, the budget cut, and the executive who asked whether this could be “leveraged”
The old “stochastic parrot” criticism of language models gets mangled into AI is a dumb autocomplete machine that can do nothing. That is too simple. Models can find real patterns and make genuinely useful suggestions.
The sharper warning is this: fluent language cannot prove that the inherited map still matches the territory.

The terrain neglected to receive the update.
A system can flag a likely range shift in a forest, a respiratory cluster, or a promising molecule. Great. Then comes the small matter of finding out whether the thing is actually there.
Somebody must collect the sample. Calibrate the instrument. Get consent. Check the patient. Document where the data came from. Determine whether the pattern is meaningful or merely a software-induced Rorschach blot.
AI can generate a hypothesis in seconds.
The universe still requires follow-up.
“Ground truth” is not a magic rock, either
The phrase ground truth has the pleasant sound of a giant granite fact rolling downhill to crush ambiguity. In practice, it usually means a label made by specific people, using specific instruments, definitions, and institutional habits—any of which may be incomplete or wrong.
So no, the alternative to AI is not some sacred human gaze, free from error and funded by moonbeams.
The goal is better evidence: transparent collection, independent checking, repeat observation, and enough humility to admit what was not measured.
That is where situated knowledge matters. Not because locals are mystical forest wizards, but because they know things the remote system often does not:
-
A grower knows the field is dry because a creek was diverted upstream.
-
A resident knows the satellite map missed an impassable road or mislabeled an informal settlement.
-
A clinician knows the tidy diagnostic category does not fit the actual person coughing in front of them.
-
A community member may recognize an event, language, or risk that never made it into a major dataset in the first place.
The model sees a record. Someone on the ground sees why the record is lying.

Excellent record. Shame about the thing it recorded.
Citizen science is useful—until it becomes Pokémon Go for free labor
Platforms such as iNaturalist and eBird have helped researchers track species distributions, migration, invasive organisms, and conservation needs. Public observations can be enormously valuable.
They also reveal the charming flaw in the idea that more data automatically means better data: people tend to record wildlife where people already are.
Observations cluster around roads, cities, tourist sites, reliable internet, affluent users, and charismatic animals. The database gets a thousand photos of a hawk and approximately three suspiciously blurry records of whatever small, ugly organism is quietly collapsing in a wetland no one can reach.
| What the platform says | What the dataset may actually contain |
|---|---|
| “A global biodiversity record” | Places users could access and species they noticed |
| “Community-powered science” | Useful but uneven observations needing expert review |
| “More data” | More data about what was already visible |
This is not a failure of participants. They are doing what they can with time, phones, transport, jobs, childcare, physical limits, and a life that does not include unlimited Tuesday afternoons hunting obscure lichens.
It is a failure of the fantasy that unpaid enthusiasm can substitute for public infrastructure.

The reward is a badge, exposure, and maybe trench foot.
Public health has no “skip ad” button
The same problem becomes less cute when the subject is disease.
COVID-19 showed what genomic and wastewater surveillance can do: reveal outbreaks and changing patterns before everybody has fully settled into denial. It also showed how unevenly that capacity is distributed.
An algorithm can spot a worrying signal in the data it receives. It cannot manufacture:
-
Representative surveillance
-
Safe, fast sample collection
-
Functioning diagnostic laboratories
-
Trust from people asked to share health information
-
Local health departments that have not been hollowed out and told to “innovate”
What is absent from a dataset is often absent for reasons with names: poverty, geography, language, race, political power, lack of access, disease severity. Scale does not cure that. Scaling a biased pipeline can simply produce industrial quantities of blindness.

Now in bulk, with improved throughput.
Robots can visit Mars. Your cluttered garage remains a frontier.
Machines will expand the senses of science: drones, satellites, camera traps, lidar, low-cost sensors, autonomous vehicles, laboratory robots. Self-driving labs can run promising chemistry and materials experiments. Machine learning can sift giant astronomy surveys while people inspect the weird edge cases—the things that do not resemble the objects the system was trained to recognize.
This is good. Use the machines.
But physical work has the insulting habit of occurring in physical places.
A robot outside a controlled setting has to deal with irregular ground, weather, contamination, fragile material, low power, dead connectivity, maintenance, permissions, farms, clinics, homes, protected landscapes, and the ancient curse known as “someone moved the thing.”
Cloud intelligence scales beautifully because it gets to live in a climate-controlled data center. Field intelligence has mud on it.
Local context is not a bug awaiting a bigger model. It is the work.

Mars was easier. Nobody stores rakes on Mars.
Synthetic data cannot discover what nobody bothered to observe
Generated text, images, and data can help with simulation and controlled training. Fine. But synthetic material inherits the assumptions of whatever produced it. It cannot independently reveal a fact nobody has measured.
If the information environment fills with machine-made imitations of previous information, provenance becomes more important, not less. Otherwise we get the intellectual equivalent of feeding a photocopier its own copies until the office gradually forgets what a face looks like.
The same goes for grand talk of AI “discovering science” by itself. AI can propose conjectures, representations, algorithms, and experimental paths that humans might miss. Formal systems can verify proofs under stated axioms.
But an empirical claim still has to survive contact with the empirical world. A beautiful theory does not become medicine, climate science, or materials engineering because it received a standing ovation from a GPU.
The real question is who gets paid, and who gets mined
There is a pleasant version of this future: communities help define research questions, document the places they know, govern their data, and share in the benefits. Paid local science. Better public-health surveillance. Public laboratories. Biodiversity monitoring. Useful tools in more hands.
And there is the version we are already very good at building: companies harvest location traces, health records, photos, worker observations, and Indigenous ecological knowledge, then call it participation because “extraction” tested poorly in branding meetings.
The photograph of a manzanita may look identical in either database.
The social arrangement behind it is not.
| The sales pitch | Plain English |
|---|---|
| “Help improve the model” | Give an owner elsewhere another free input |
| “Citizen science” | Valuable work, possibly unpaid |
| “Data-driven innovation” | Who controls the data, and who receives the gains? |
Consent, data minimization, privacy, security, community governance, benefit sharing, and data trusts or cooperatives are not bureaucratic garnish. They are the difference between collaboration and a more polite kind of taking. Indigenous data governance principles such as CARE—collective benefit, authority to control, responsibility, ethics—make the obvious point that the technology sector keeps rediscovering as if it were ancient wisdom:
Useful does not mean yours.
Stop measuring humans by their ability to imitate office software
The fear of replacement is not a quaint objection to progress. When work determines income, health insurance, status, and the ability to remain indoors, automation does not feel like liberation from drudgery. It feels like somebody offering to “free” you from rent-paying.
If AI makes routine symbolic work abundant, the answer cannot be to shove people into unpaid monitoring, tagging, and “community contribution” while owners of models, compute, and data collect the productivity gains.
Human value is not typing faster than a machine that has eaten the internet. The work that matters may be harder, messier, and less scalable: noticing what the model missed, testing its claims, maintaining the systems that gather evidence, protecting people whose data is being requested, and deciding which questions deserve asking.
The chatbot can describe the manzanita beautifully.
Somebody still has to notice when it is dying.

Fortunately, the brochure remains in excellent health.