When humans are needed
Where real capture beats synthetic data — and where it does not.
Harbor captures first-person video from real work. Labs still use synthetic data. This page is the split: when you need people, and when you do not.
Use people when you need to beat today’s best model
You need humans when you want to:
- Push a capability past what the best model can already do
- Fine-tune a skill the model cannot do today
- Measure an AI system against real work, not a lab demo
Web scrapes and synthetic footage miss real work: the shop floor, the recovery, the messy hour that separates a demo from a system you can deploy.
Harbor’s focus is audio-visual AI — vision, speech, wearables, robotics, and agentic workflows — because those are the places synthetic shortcuts break first.
Use synthetic data when you are catching up
You may not need a full human capture programme when:
- Prompt changes still move the metric
- You only need a solid baseline close to a frontier model
- You are distilling a stronger model into a smaller one for a known task
Synthetic data copies a teacher model. It also copies that model’s errors. It is useful for catching up. It is a poor way to invent a new capability.
Where synthetic data fails
When you generate training data from a model, you get that model’s style, bias, and blind spots. Models also prefer their own answers over human ones. That loop does not push the frontier.
Agents that plan and act in the world are a clear example. Today’s frontier models are not good enough at breaking down a real task. Sampling their own plans, without human review, is not a dataset you can trust.
How Harbor uses AI with humans
AI still belongs in the pipeline. Harbor uses models to triage, flag errors, and sanity-check work. People remain the source of new signal: consented capture, Passport identity, and specialist review against a published rubric.
Need a programme scoped? Talk to Harbor.