API to livestream human preferences for training and evaluating multimodal models

Get 5000+ human preferences per minute to enable online RL, refine model behaviour and evaluate outputs.

Trusted by folks at
IdeogramThinking MachinesStanford UniversityVanderbilt University
API

One request, 100M+ people,
200+ countries

Train models, no call required. Get a key, send your first request, and see real answers come back.

Task type

Modality

“Rate this image on a scale of 5”

curl -X POST https://api.trydatapoint.com/data-labelling/v1/jobs \
  -H "X-API-Key: $DATAPOINT_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "vqa-charts-v3",
    "instruction": "Rate this image on a scale of 5",
    "task_type": "rating",
    "response_options": { "scale": [1, 2, 3, 4, 5] },
    "max_responses_per_datapoint": 5,
    "datapoints": [{
      "media": { "subject": [{ "url": "dp://b6a1cd3f/image_a.png", "type": "image" }] }
    }]
  }'
POST/v1/mediaPOST/v1/jobsGET/v1/jobs/:id/results
Read the docs
Quality

Every answer earns its place in your report.

Every response goes through multiple checks before it reaches you — identity, behavior, attention, bias. Most platforms count responses. We disqualify them.

Read docs
A funnel of stacked checks narrowing toward your report
Trust score71/100+16 after attention
Attention

Embedded attention checks catch anyone skimming instead of reading.

Use cases

One live stream,
every stage of the model

Wherever a human judgment should move the model — inside the training loop, between checkpoints or after launch — the same API delivers it.

01

Faster training

Swap week-long labelling rounds for a live stream. Preferences arrive as fast as your loop can use them.

5,000+/ minlive streambatch labelling
02

Online RL

Pipe human rewards straight into the policy update — per step, not per batch.

03

Continuous benchmarking

Re-run your eval suite on every build and watch the ratings move nightly.

04

Hill-climbing specific abilities

Target one skill — text rendering, instruction following, tone — and climb it with focused preference data.

05

Evaluating internal checkpoints

Pit checkpoints head-to-head before anything ships, with win rates and rationales for every pair.

06

Post-training for a specific demographic

Recruit raters by locale, language, age or profession, so the model learns the taste of the people who will actually use it.

200+countries
DP-Bench

Ranked by eligible
human votes

Elo ratings computed from pairwise preference votes that cleared every screen. The interval is what matters — models whose intervals overlap are not separated by this evidence.

#ModelElo95% interval
1
2
3
4
5
6

Ranked by eligible human preference votes

View all leaderboards
Contact

Tell us what you need judged

A research engineer answers — not a sales queue. Send the modality and the monthly volume; a panel spec and a price come back within a working day.