Chosen and rejected, per pair
Every rollout pair comes back with the rater’s pick, ready to use as a reward signal or as a DPO pair.
Rollouts go out through the Datapoint API, real people judge them, and rewards stream back while the policy is still training. Low-trust answers are dropped before they ever reach the update.
By the time a labelling round comes back, the policy that produced those samples is gone. Online RL needs feedback on what the model is doing now, at the speed it learns.
BatchDays to weeks per labelling round
Seconds per judgment, streamed as they land
BatchOutputs of a checkpoint the policy has already moved past
The rollouts your policy produced this step, so feedback stays on-policy
BatchOne reward-model refresh per batch
Human reward on every policy update
BatchFound after the run, in the eval
Caught while it trains, by people a proxy reward cannot stand in for
BatchA fixed annotator workforce
Everyday people, targeted by locale, language, age or profession
Each preference arrives as structured data your training loop can use directly, with the quality work already done.
Every rollout pair comes back with the rater’s pick, ready to use as a reward signal or as a DPO pair.
Raters can write why they chose. Use it to debug the policy or to train a reward model that reasons.
Attention checks, response-time floors, agreement scoring and device verification run before a preference reaches your update.
Per-datapoint agreement stats tell you which preferences are clear-cut and which are a coin flip, so you can weight them.
Compare, rate or rank image, audio and video rollouts on the same API, and mix them within one run.
Scale the stream up for a sweep and down between runs. No panel to manage, no workforce to contract.