Introducing Streams

Online RL with live human feedback for multimodal models

Send image, audio and video rollouts while your model trains. Real people from a pool of 1B+ raters judge them in real time, and their preferences come back as a continuous reward stream.

How it works

One loop. Live human feedback.

Rollouts go out through the Datapoint API, real people judge them, and rewards stream back while the policy is still training. Low-trust answers are dropped before they ever reach the update.

The Streams loop: your model sends rollouts to the Datapoint API, real people worldwide judge them, and rewards stream back while it trains, with low-trust answers dropped before the reward model updates.
Reward model
v8updated
reward0.84
Rollouts →← Preferences
DatapointAPI
840streamed · 10,000+/min
Real people, worldwide
1B+ raters · 200+ countries
  1. 01Your model generates rollouts
  2. 02Real people judge them
  3. 03Rewards stream back while it trains
  4. 04The reward model updates, and keeps going
10,000+human preferences per minute
1B+raters on tap
200+countries represented
Why online

Batch labelling trains on last month’s model

By the time a labelling round comes back, the policy that produced those samples is gone. Online RL needs feedback on what the model is doing now, at the speed it learns.

Feedback latency

BatchDays to weeks per labelling round

Seconds per judgment, streamed as they land

Which samples get judged

BatchOutputs of a checkpoint the policy has already moved past

The rollouts your policy produced this step, so feedback stays on-policy

Update granularity

BatchOne reward-model refresh per batch

Human reward on every policy update

Reward hacking

BatchFound after the run, in the eval

Caught while it trains, by people a proxy reward cannot stand in for

Who judges

BatchA fixed annotator workforce

Everyday people, targeted by locale, language, age or profession

What comes back

More than a thumbs up

Each preference arrives as structured data your training loop can use directly, with the quality work already done.

01

Chosen and rejected, per pair

Every rollout pair comes back with the rater’s pick, ready to use as a reward signal or as a DPO pair.

02

Rationales, not just clicks

Raters can write why they chose. Use it to debug the policy or to train a reward model that reasons.

03

Low-trust answers dropped in flight

Attention checks, response-time floors, agreement scoring and device verification run before a preference reaches your update.

04

Agreement on every datapoint

Per-datapoint agreement stats tell you which preferences are clear-cut and which are a coin flip, so you can weight them.

05

Images, audio & video

Compare, rate or rank image, audio and video rollouts on the same API, and mix them within one run.

06

Elastic, pay per response

Scale the stream up for a sweep and down between runs. No panel to manage, no workforce to contract.

Get started

Put real people in your training loop

Tell us your modality, your step rate and the judgment you need. A research engineer comes back with a panel spec and a price within a working day.