# Streams: online RL with live human feedback

> Send image, audio and video rollouts while your model trains. Real people from a pool of
> 1B+ raters judge them in real time, and their preferences come back as a continuous
> reward stream.

Online reinforcement learning needs feedback on what the policy is doing now, at the speed
it learns. Streams turns the Datapoint API into that feedback channel: rollouts go out,
real people judge them, and rewards stream back while the policy is still training.
Low-trust answers are dropped before they ever reach the update.

## How it works

1. **Your model generates rollouts.** Your policy produces candidate outputs for the
   current step (images, audio or video) and posts them as comparison, rating or ranking
   jobs.
2. **Real people judge them.** Everyday respondents recruited through a global in-app ad
   network answer within seconds.
3. **Rewards stream back while it trains.** Each judgment returns as structured data with
   the quality work already done. Attention checks, response-time floors, agreement
   scoring and device verification run in flight.
4. **The reward model updates, and keeps going.** Human reward lands on every policy update,
   so feedback stays on-policy instead of describing a checkpoint the model has already
   moved past.

## Streams in numbers

- 10,000+ human preferences per minute
- 1B+ raters on tap
- 200+ countries represented

## Batch labelling vs. online with Datapoint

|                          | Batch labelling                                           | Online with Datapoint                                                    |
| ------------------------ | --------------------------------------------------------- | ------------------------------------------------------------------------ |
| Feedback latency         | Days to weeks per labelling round                         | Seconds per judgment, streamed as they land                              |
| Which samples get judged | Outputs of a checkpoint the policy has already moved past | The rollouts your policy produced this step, so feedback stays on-policy |
| Update granularity       | One reward-model refresh per batch                        | Human reward on every policy update                                      |
| Reward hacking           | Found after the run, in the eval                          | Caught while it trains, by people a proxy reward cannot stand in for     |
| Who judges               | A fixed annotator workforce                               | Everyday people, targeted by locale, language, age or profession         |

By the time a labelling round comes back, the policy that produced those samples is gone.
Online RL needs feedback on what the model is doing now.

## What comes back

Each preference arrives as structured data your training loop can use directly:

- **Chosen and rejected, per pair.** Every rollout pair comes back with the rater's pick,
  ready to use as a reward signal or as a DPO pair.
- **Rationales, not just clicks.** Raters can write why they chose. Use it to debug the
  policy or to train a reward model that reasons.
- **Low-trust answers dropped in flight.** Attention checks, response-time floors,
  agreement scoring and device verification run before a preference reaches your update.
- **Agreement on every datapoint.** Per-datapoint agreement stats tell you which
  preferences are clear-cut and which are a coin flip, so you can weight them.
- **Images, audio & video.** Compare, rate or rank image, audio and video rollouts on the
  same API, and mix them within one run.
- **Elastic, pay per response.** Scale the stream up for a sweep and down between runs.
  No panel to manage, no workforce to contract.

## Get started

Tell us your modality, your step rate and the judgment you need. A research engineer comes
back with a panel spec and a price within a working day.

- Book a demo: https://calendly.com/akshat-trydatapoint/30min
- API docs: https://trydatapoint.com/docs
- Human preference data API: https://trydatapoint.com/human-preference-data-api
- Glossary (RLHF, DPO, reward models): https://trydatapoint.com/glossary
