2026-09-21

BEHAVIOR-1K 2026 data composition

This page describes the 100-task benchmark from the files the evaluator reads. Those files are its constants, the robot configuration it loads, the goal definitions it scores against, and all 20,000 recorded demonstrations. The facts come first. The four places where the challenge documentation disagrees with the code come after. Two of them fall on the 50 tasks added for 2026, which nobody has run.

100 tasks20,000 demos1,018 BDDL definitions23 submissions from 2025
The data

What the benchmark is

These facts come before any of the questions that follow. Every figure here comes from the evaluator's own constants, the robot configuration it loads, or the recorded demonstrations. None comes from the challenge documentation.

The corpus

100
tasks
50 carried from 2025, 50 new, none dropped
20,000
teleop demos
exactly 200 per task, no exceptions
341s
median demo
10,233 steps at 30 Hz
7
scenes
4 new, hosting 24 tasks exclusively
0.089
2025 field-mean Q
unreported tasks counted as 0; 0.236 over reported ones
5
tasks at zero
183 rollouts run, none scored above 0
The new half is shorter, not longer. Carried tasks run 397 s on average. The 50 new tasks run 306 s, which is 23 % less. Length is one of the stronger difficulty predictors, so the 2026 set may be easier per task than the 2025 set. That is the opposite of the usual direction for a second-year benchmark.
Ten of the 20,000 episodes run under 40 seconds. The shortest is 148 steps, or 4.9 s, in putting_away_Halloween_decorations, where the median episode runs 7.7 minutes. Its six video streams span 4.93 s and match the recorded length exactly, so the recording is short rather than truncated. Ten episodes in 20,000 move no statistic. Filter on length if you sample episodes individually.

What the robot is, sees and does

RobotR1Pro, a bimanual mobile manipulatoreval/r1pro.yaml
Control rate30 Hz actions over 120 Hz physics — one action every 33 mseval_utils.py:177,179
Camerasthree — head at 720×720, each wrist at 480×480, RGB and deptheval_utils.py:10,11
Proprioceptionbase velocity, per-arm joint position and velocity, end-effector pose, gripper and trunk stater1pro.yaml proprio_obs
Action23 dimensions: base 0:3, torso 3:7, left arm 7:14, left gripper 14, right arm 15:22, right gripper 22eval_utils.py ACTION_QPOS_INDICES
Control modearms and trunk take absolute joint positions; the base takes a velocity, capped at 0.75 m/s and 1.0 rad/sr1pro.yaml controller_config
Graspingassisted — an object attaches when the gripper closes near itr1pro.yaml grasping_mode

How it is scored

Q-scorelen(satisfied) / (len(satisfied) + len(unsatisfied)) over the goal predicates, read at episode endbehavior_task.py:410
Goal predicates1 to 12 per task, median 3; 66 of 100 goal blocks contain a quantifier that expands at runtimeproblem0.bddl
Episode timeoutint(length × 1.5) — from 3,224 steps on the shortest task to 39,090 on the longesteval_utils.py:13, evaluator.py:163
Test instances40 per task, ids 301–340: the first 20 public, the last 20 hiddeneval_utils.py:14–17
Tiebreakerssimulated time, base distance and end-effector displacement, each normalised as human ÷ robot, so above 1 beats the human demonstrationtask.jsonl, score_utils.py
One consequence to carry into everything below. Q is a ratio whose denominator is the task's goal-predicate count, and that count is often 1. Partial credit is available on paper across the benchmark, and absent in practice on a quarter of it.
Data fit

How well the corpus suits each candidate

The corpus is fixed. These three methods need different things from it, and it serves them unequally. fits means the data suits the method as it is. workable means usable at a cost. outside means outside what the method has been shown to handle.

What the data suppliesverified from code and demos Diffusion PolicyDDPM, 278M, per-task π0.5 baseflow matching, ~3B, multi-task Comet pt50π0.5 fine-tuned on 2025
200 demos per task, every task fitsexactly its own setting — 200 proficient-human demos on every Robomimic task fits20,000 total across tasks fitstrained on range(200), all of them
Median episode 10,264 steps
shortest task 2,150
outsideits longest benchmark caps at 700 steps. The shortest BEHAVIOR task is 3.1× that; the median is 14.7×, and 37× Push-T workabledesigned for long horizons, but its own reported evaluations are far shorter workabledemonstrated at this length, at Q 0.25 held-out
23-dimension action
mixed position and velocity
outsideits benchmarks run 2–14 dims, all single-mode control fitsmodel action dim is 32; 23 is zero-padded to fit fitssame
979,200 px rendered per step
720² head + 2×480² wrists, RGB and depth
outsidetrained at 84²–240²; the render is 69× its Square setting workablethe serving wrapper resizes RGB to 224² and keeps depth at 720, so most of the render is discarded before the model sees it workablesame wrapper
Per-segment skill language
annotations/, all 100 tasks; but no one-sentence instruction for the new 50
workabledoes not use language, so neither the annotations nor the missing instruction change anything for itfitsdense phase text is exactly what a language-conditioned model can use — if the loader reads itworkablesame, and openpi’s loader looks for an orchestrators/ directory the release does not publish
100 tasks outsideone policy per task in every published result — 100 training runs fitsmulti-task by construction workabletrained on the 50 carried tasks; has never seen the 50 new ones or the 4 new scenes
210.9M frames, 1,953 hours outsideone BEHAVIOR task alone is 2.05M frames, 26× the whole Square dataset workablewithin scale, but 3.27 TB gates who can train on it workabletrained on the 2025 subset at batch 256 for 50k steps
30 Hz control workableits real-robot setting is 10 Hz; the rate behind its simulated ablations is not stated in the paper fitschunk 32 spans 1.07 s at this rate fitschunk 32, matching
Diffusion Policy fits the corpus least well, and is the most useful control. Five of the eight rows fall outside what it has been shown to handle. The one row that matches exactly is 200 demonstrations per task. It is also the only one of the three that does not use language, so the missing instructions cost it nothing. It is small enough to train per task on one workstation GPU. That makes it a poor submission candidate and a useful control: it differs from the published Diffusion Policy results in task regime only, and from the π0.5 arm in architecture only.
Comet pt50 fits the corpus best, and is the most constrained. Its training set was behavior-1k/2025-challenge-demos. The 50 tasks it saw carry into 2026 with their demonstration lengths unchanged. Not one carried task moved. Half the 2026 benchmark is therefore in distribution for it. The other half is not, and neither are the four scenes that hold only new tasks.
The resolution result has a counter-intuitive cause. The evaluator renders 979,200 RGB pixels per step, and the serving wrapper resizes them to 224². Extra detail reaching the policy therefore cannot explain the measured benefit of full-resolution rendering. The model sees the same 224² image either way. The likely mechanism is downsampling quality. A 720² render reduced to 224² is better antialiased than a native 224² render. This is testable, and nobody has tested it.
What looks wrong

Four things that do not behave as documented

Ordered by how much each one should change a decision. Every claim below cites the file it came from. None rests on a docstring or a README.

  1. Task durations do not match the demos, for the new 50 only
    All 50 carried-over tasks agree with the evaluator's demo length to within a second. Of the 50 new ones, two do. Median error 13.3 %, worst +31.6 % on cook_brussels_sprouts. Scoring is unaffected. evaluator.py:163 reads length from task.jsonl and never reads this field. But any plan built on duration is wrong on the half of the benchmark that nobody has run. The table below gives each task.
    source task_data.json vs datasets/2026-challenge-task-instances/metadata/task.jsonl · evaluator.py:163
  2. 24 tasks have exactly one goal predicate, so Q is binary
    Q-score is len(satisfied)/(len(satisfied)+len(unsatisfied)). A one-predicate task has no partial credit. Q is 0 or 1. This includes turning_on_radio, the task that every published baseline is tuned on. The median task across the benchmark has 3 goal predicates.
    source behavior_task.py:410 · bddl3/bddl/activity_definitions/*/problem0.bddl
  3. The new 50 tasks have no high-level instruction, though the dataset is richly annotated
    All 50 carried-over tasks carry an instruction string in task_data.json. None of the 50 new tasks carry one. meta/tasks.parquet holds only a task index, and the per-episode tasks field holds the bare snake_case name. That is 100 distinct values across all 20,000 episodes, not 20,000 annotations. The dataset does supply language, in a directory that none of those three files point to. annotations/ is the largest directory in the release, at 20,002 files, one JSON per episode for all 100 tasks. Each file carries per-segment skill text with frame ranges. For vacuuming_floors, one of the new 50: "move to" [0,416] on ["vacuum"], then "pick up from" [416,592] on ["vacuum","floors"], each tagged with a skill type. The gap is therefore the one-sentence task instruction, not language in general. An earlier version of this page stated that no language existed for the new 50 tasks. That was wrong. The check covered only meta/.
    source task_data.json · meta/tasks.parquet · meta/episodes/chunk-*/file-*.parquet · annotations/task-0069/episode_00690030.json · HF list_repo_files
  4. The five tasks nobody scored on in 2025 share one predicate
    wash_dog_toys, clean_a_patio, clean_a_trumpet, cook_cabbage and make_pizza hold Q = 0.000 in all 23 submissions, and not one rollout of them ever scored above zero. Far fewer teams tried them than that makes it sound. Each was reported by only 3 to 6 of the 23 submissions, because most submissions are partial and an unreported task is recorded as 0.000. What survives is still a clean result: across the 183 rollouts that were actually run on these five tasks, by the two teams that reported the whole benchmark and four others, every single one scored 0.000. The privileged track, whose submissions may read ground-truth simulator state, is thinner evidence than it appears here: one privileged submission reported two of the five, both at zero, and none reported the other three. Three of the five are scored on covered, which appears in 60 % of the zero-score tasks against 5 % of the solved ones. covered is a particle-system state, such as dust, water and stains. The field did not fail at long horizons on these tasks. It failed at fluids.
    source docs/challenge_submissions/*.json · goal predicates from problem0.bddl
Finding 01, in full

Duration, task by task

stated is task_data.json:duration, an integer field. actual is task.jsonl:length ÷ 30 Hz, the figure the evaluator uses. Because stated holds whole seconds, it cannot be more precise than ±1 s. Rows inside that band are therefore marked as agreeing, rather than shown a rounding artefact. timeout is the step cap the evaluator derives, int(length × 1.5). Sorted by disagreement.

50 / 50
carried tasks agree
within one second, every one
2 / 50
new tasks agree
halve_an_egg, unloading_the_car
26.6s
median new-task error
13.3 % of the true length
27 / 21
longer / shorter
than the website states; 52 agree
Taskid used by --task-nameSetadded 2026, or carriedStatedtask_data.json durationActualtask.jsonl length÷30Δ show far the real length differs — ▲ longer, ▼ shorter, ● agreesΔ %share of true lengthTimeoutstep cap, length×1.5
cook_brussels_sprouts687522.1▼−164.9−31.6%23,492
setting_the_table727593.6▼−133.4−22.5%26,712
tidying_bathroom541433.4▼−107.6−24.8%19,504
tidying_living_room316418.9▲+102.9+24.6%18,848
setup_a_bar_for_a_cocktail_party350447.3▲+97.3+21.8%20,128
rearrange_your_room515428.5▼−86.5−20.2%19,282
stacking_wood447531.3▲+84.3+15.9%23,909
freeze_fruit339420.6▲+81.6+19.4%18,928
make_gift_bags_for_baby_showers242321.7▲+79.7+24.8%14,476
thawing_frozen_food348274.3▼−73.7−26.9%12,344
put_together_a_basic_pruning_kit448375.5▼−72.5−19.3%16,896
store_batteries328256.8▼−71.2−27.7%11,556
putting_dirty_dishes_in_sink337405.0▲+68.0+16.8%18,222
dispose_of_glass240306.5▲+66.5+21.7%13,792
putting_away_toys438376.6▼−61.4−16.3%16,945
organizing_school_stuff426486.4▲+60.4+12.4%21,887
sorting_books_on_shelf200258.0▲+58.0+22.5%11,609
bringing_paper_to_recycling318373.8▲+55.8+14.9%16,821
laying_tile_floors520478.4▼−41.6−8.7%21,526
make_rose_centerpieces111151.1▲+40.1+26.5%6,798
dispose_of_batteries441480.9▲+39.9+8.3%21,642
re_shelving_library_books281317.6▲+36.6+11.5%14,292
cook_a_frozen_pie323288.9▼−34.1−11.8%13,001
store_honey193225.6▲+32.6+14.5%10,150
sorting_bottles_cans_and_paper312342.1▲+30.1+8.8%15,395
sweeping_garage126149.1▲+23.1+15.5%6,709
organizing_art_supplies228206.5▼−21.5−10.4%9,291
clean_your_rusty_garden_tools484505.3▲+21.3+4.2%22,736
installing_a_scanner122142.0▲+20.0+14.1%6,387
scrubbing_bathroom_floor123105.1▼−17.9−17.0%4,730
installing_smoke_detectors6885.6▲+17.6+20.6%3,853
installing_a_modem6380.4▲+17.4+21.6%3,619
cook_broccolini137153.9▲+16.9+11.0%6,926
boxing_food_after_dinner240223.3▼−16.7−7.5%10,046
cleaning_up_branches_and_twigs434417.6▼−16.4−3.9%18,790
vacuuming_floors6580.4▲+15.4+19.2%3,617
collecting_aluminum_cans325340.1▲+15.1+4.4%15,304
clean_a_keyboard142129.2▼−12.8−9.9%5,814
installing_a_fax_machine148135.9▼−12.1−8.9%6,115
carrying_out_garden_furniture358368.6▲+10.6+2.9%16,588
cook_a_brisket251244.5▼−6.5−2.7%11,002
composting_waste115121.3▲+6.3+5.2%5,457
make_cabinet_doors120114.8▼−5.2−4.5%5,163
polishing_shoes363357.9▼−5.1−1.4%16,106
packing_meal_for_delivery300295.3▼−4.7−1.6%13,289
clean_up_broken_glass286289.7▲+3.7+1.3%13,037
turning_out_all_lights_before_sleep337333.5▼−3.5−1.0%15,009
store_produce316317.4▲+1.4+0.4%14,283
setting_mousetraps339339.9●——15,294
unloading_the_car378377.4●——16,985
halve_an_egg212212.6●——9,565
putting_shoes_on_rack258257.5●——11,589
clean_boxing_gloves275274.5●——12,353
cleaning_up_plates_and_food457456.5●——20,544
cook_cabbage471471.5●——21,217
moving_boxes_to_storage487486.5●——21,894
collecting_childrens_toys640639.5●——28,779
hanging_pictures8079.6●——3,581
attach_a_camera_to_a_tripod130130.4●——5,867
picking_up_trash176175.6●——7,901
bringing_water315314.6●——14,157
outfit_a_basic_toolbox355354.6●——15,956
chopping_wood358358.4●——16,127
clean_a_patio402402.4●——18,106
clearing_food_from_table_into_fridge436435.6●——19,602
putting_away_Halloween_decorations460459.6●——20,682
getting_organized_for_work522522.4●——23,506
make_pizza640639.6●——28,780
boxing_books_up_for_storage808807.6●——36,341
chop_an_onion213213.3●——9,599
wash_a_baseball_cap278278.3●——12,524
putting_up_Christmas_decorations_inside457457.3●——20,578
turning_on_radio7271.7●——3,224
picking_up_toys630629.7●——28,335
storing_food662662.3●——29,803
assembling_gift_baskets869868.7●——39,090
canning_food766765.8●——34,463
preparing_lunch_box275274.8●——12,367
spraying_fruit_trees278278.2●——12,518
cook_hot_dogs305304.8●——13,717
sorting_vegetables397396.8●——17,855
freeze_pies415415.2●——18,682
bringing_in_wood451451.2●——20,303
carrying_in_groceries476475.8●——21,412
slicing_vegetables495494.8●——22,267
rearranging_kitchen_furniture298298.1●——13,414
setting_the_fire304303.9●——13,677
putting_dishes_away_after_cleaning365365.1●——16,430
tidying_bedroom368367.9●——16,556
wash_dog_toys374374.1●——16,834
can_meat395394.9●——17,770
sorting_household_items527526.9●——23,711
loading_the_car641640.9●——28,839
clean_up_your_desk714713.9●——32,126
make_microwave_popcorn108107.9●——4,856
clean_a_trumpet177176.9●——7,960
set_up_a_coffee_station_in_your_kitchen209208.9●——9,399
spraying_for_bugs216216.0●——9,719
hiding_Easter_eggs254254.0●——11,429
cook_bacon256256.0●——11,519
The timeout column is the one that matters. The evaluator derives it from length rather than duration, so every episode receives the correct budget. This field damages planning, not scoring. The damage falls entirely on the 50 tasks with the least available information.
The metric

Where Q-score points actually live

Q is the fraction of goal predicates satisfied at episode end, averaged over tasks. The predicate structure of each task therefore is the scoring surface. That surface is shallower than it appears. 24 tasks are all-or-nothing.

1
fewest predicates
24 tasks; Q is 0 or 1
3
median predicates
one predicate = 33% of Q
12
most predicates
setup_a_bar_for_a_cocktail_party
66
tasks with quantifiers
true denominator set at runtime
Goal predicateBDDL state checked at episode endTasksof 100 that use itShare
inside54
ontop40
open23
nextto16
covered13
real9
toggled_on7
cooked6
attached4
contains4
Two predicates carry the benchmark. inside and ontop appear in 54 and 40 tasks. A policy that places one object inside or on top of another reaches more of the scoring surface than any other single skill. covered appears in only 13 tasks, and the entire 2025 field scored zero on it.
Progress

Which predicates pay out, and how far anyone gets

Each of the 23 submissions (listed under Provenance) reports a Q for every individual rollout. Q is satisfied ÷ total. A value on an N-predicate task therefore decodes to the number of predicates that rollout achieved. 3,256 rollouts decode this way.

Getting some credit versus finishing

This chart covers the 21 tasks whose goal uses a single predicate type, so achievement attributes to that type without ambiguity. The bar runs from the completion rate to the any-credit rate. A long bar means policies start the job often and finish it rarely.

inside
14.2% → 66.7%
ontop
37.7% → 66.4%
cooked
64.0% → 66.0%
toggled_on
57.2% → 57.2%
real
23.2% → 55.4%
attached
18.2% → 18.2%
covered
9.9% → 17.9%
completed every predicatesatisfied at least onebar between the two = the partial-credit zone

How deep into the chain rollouts get

All scored tasks, grouped by how many goal predicates they have. Each row gives the share of rollouts that satisfied 0, 1, 2 and so on up to N. Darker means further along. The grey block on the left is complete failure.

N = 1
68.0%32.0%
918
N = 2
54.7%12.1%33.2%
404
N = 3
47.9%13.4%20.8%17.8%
365
N = 4
45.2%28.1%
922
N = 5
43.2%20.9%18.9%12.2%
148
N = 6
22.4%20.3%14.6%13.0%15.1%
192
N = 7
52.5%
160
N = 8
40.7%17.4%16.3%15.1%
86
0 predicatesfirstmiddleall Nright-hand figure is the rollout count
inside supplies the partial credit. ontop is where tasks finish. Both earn something on about two thirds of rollouts. But ontop completes on 37.7 % against 14.2 % for inside. The field can place objects on things. It can start to place them in things, and seldom finishes.
covered and attached rarely start. Fewer than 19 % of rollouts earn any credit on either. For attached, the any-credit rate and the completion rate are the same number. It never succeeds partly. These two are the particle and fastening states, and they set the floor of the benchmark.
What this cannot tell you. Q gives the count of satisfied predicates, never which ones. On a mixed-type task, 1 of 4 does not name which of the four. Satisfaction order needs per-frame replay. A replay harness of this kind already does this, recording satisfied and unsatisfied predicate indices per frame and timestamping each change. It has run on two tasks only. Extending it across the benchmark is a GPU job, and it is the only way to find which predicate comes first.

The same numbers as a table

Predicatesingle-type tasks only TasksRollouts Any creditCompleted Mean Q
inside747266.7%14.2%0.338
ontop222366.4%37.7%0.508
cooked15066.0%64.0%0.65
toggled_on118057.2%57.2%0.572
real15655.4%23.2%0.371
attached213218.2%18.2%0.182
covered737417.9%9.9%0.139
Difficulty

What predicts a hard task

Rank correlation against the 2025 field-mean Q, over the 50 carried tasks. Those are the only tasks with measured outcomes. All four correlations are negative. The question is which one is strongest.

Predictorknown before any rolloutSpearman ρvs 2025 field-mean QStrengthWhat it measures
end-effector displacement-0.467
how far the hands travel
base distance-0.384
how far the robot drives
demo length-0.372
how long it takes
goal predicates-0.215
how many conditions
Manipulation, not duration. The distance the hands must travel predicts difficulty better than the length of the task, and better still than the number of conditions. The tiebreaker metrics that the challenge already records are therefore the best available difficulty index. They are known in advance for all 100 tasks, including the 50 that nobody has attempted. Goal-predicate count is both the weakest predictor of the four and the least reliable, because the count is a floor on the 66 tasks whose goals are quantified. See Method.
Read every number here as a bound, not a measurement. The outcome variable is the field-mean Q, and most 2025 submissions are partial. A task that few teams reported carries mostly zeros whatever its difficulty, and the count of submissions reporting a task correlates with its field-mean Q at ρ 0.824, higher than any predictor in the table. Teams chose which tasks to run, and that choice is inside the outcome. The four predictors are still the best ordering available for the unrun 50, because they need no rollout at all.

The five tasks that never scored, on any rollout

TaskSecondsmean demoPredicatesQ denominatorGoal predicate types
clean_a_trumpet1771covered
wash_dog_toys3744covered
clean_a_patio4021covered
cook_cabbage4724contains real
make_pizza6402ontop real
Provenance

The 23 submissions this page leans on

Every 2025 figure here comes from these files, archived in the fork at docs/challenge_submissions/. Twenty-three submissions come from eighteen teams. Five teams submitted to both test sets. The privileged track could read ground-truth simulator state.

TeamAffiliationTrack Test setQoverall, 50 tasks Task SRfully solved
Robot Learning CollectiveIndependentstandardpublic0.26050.112
Robot Learning CollectiveIndependentstandardhidden0.25990.124
CometNVIDIA Researchstandardhidden0.25140.114
SimpleAI RobotBeijing Simple AI Technology Co Ltdstandardpublic0.19430.140
CometNVIDIA Researchstandardpublic0.18300.144
The North StarHuawei CRI EAI Teamstandardpublic0.17020.128
SimpleAI RobotBeijing Simple AI Technology Co Ltdstandardhidden0.15910.108
The North StarHuawei CRI EAI Teamstandardhidden0.12040.076
Embodied IntelligenceIndependentprivilegedpublic0.11100.062
Embodied IntelligenceIndependentprivilegedhidden0.09470.052
RAPPERGISTprivilegedpublic0.07500.052
tobiAlzonovastandardpublic0.07170.036
MRMRprivilegedpublic0.05120.034
RACΞLCMUstandardpublic0.01400.014
Ahri+EFFL+MLVPostech standardpublic0.01000.010
Merlin LabsIndependentstandardpublic0.00900.006
LYQRoboticsIndependentstandardpublic0.00800.008
ACTXiamen Universitystandardpublic0.00370.002
StarVLAIndependentstandardpublic0.00190.000
Cloud-DataCloud Data Technology Co Ltdstandardpublic0.00000.000
EntropyMaximumIndependentstandardpublic0.00000.000
MagikidMagikidstandardpublic0.00000.000
RobotSimArk1standardpublic0.00000.000
Most of these submissions are partial. Only 3 of the 23 report all 50 tasks; the median reports 11, and 18 report fewer than 20. Unreported tasks score 0 and stay in the denominator, which is the rule the challenge states. So the mean of 0.089 across all 50 tasks measures coverage as much as capability. Over the tasks each submission actually reports, the same mean is 0.236.
Discussion

Comet did not improve on the hidden set

Comet is the only 2025 team whose Q rose from the public set to the hidden set, by 37 %. The published files settle why. Its public submission reports rollouts for 19 of the 50 tasks. Its hidden submission reports all 50. A task with no rollouts scores 0 and stays in the denominator, so the two numbers are averages over different amounts of the benchmark.

Team Tasks reportedpublic → hidden Q-scorepublic → hidden
Comet19500.18300.2514+37%
Robot Learning Collective50500.26050.2599-0%
SimpleAI Robot15150.19430.1591-18%
The North Star13130.17020.1204-29%
Embodied Intelligence11100.11100.0947-15%
A missing task is not a task that scored zero. Both appear as 0.000 in per_task_scores, so the distinction is only visible in per_rollout_scores. A task that was run and failed keeps its ten rollout records, each holding 0.0. Robot Learning Collective's public file carries rollouts for all 50 tasks including its 9 zeros, and Comet's own hidden file carries rollouts for all 50 including its 8 zeros. The 31 tasks absent from Comet's public file carry no rollouts at all.

The same 19 tasks, measured on both sets

Cometthe 19 tasks both submissions report PublicHiddenChange
Q-score0.48160.3978-17%
Task success rate0.37890.2789-26%
Tasks improved / fell / matched4 improved, 14 fell, 1 matched
Comet fell on the hidden set, like every other team. On the tasks both submissions report, Q drops 17 % and the task success rate drops 26 %. That sits between SimpleAI Robot at −18 % and Embodied Intelligence at −15 %, measured the same way on each team's own reported set. Comet is mid-field, not an exception. The 37 % rise is the arithmetic of 31 unreported tasks entering the average, and those 31 average 0.1617 on the hidden set.
The two tasks cited as evidence of new partial credit were never reported on the public set. An earlier version of this page read collecting_childrens_toys 0.000→0.583 and storing_food 0.000→0.400 as a policy gaining partial credit. Both are in the unreported 31, and so are 23 of the 27 tasks that appear to have improved. Only four improved among the tasks reported on both sets. The rise is coverage, not behaviour.
For anyone using the released checkpoints, 0.1830 is not a policy-quality figure. It is 19 tasks of score divided by 50. Over the tasks that submission actually reports, the public figure is 0.4816. A reproduction that runs all 50 public tasks is measuring something neither number measures, so comparing it against 0.1830 or against the hidden-set 0.2514 will mislead in opposite directions. Robot Learning Collective took the top hidden-set score at 0.2599, from a submission that reports all 50 tasks on both sets.
Method

How the predicates were counted, and why the count is a floor

Goal predicates were counted by parsing the (:goal …) block of each problem0.bddl and taking the leaf predicates. For a goal like turning_on_radio's, that is exact:

(:goal (and (toggled_on ?radio_receiver.n.01_1)))

One predicate written, one predicate scored. But 66 of the 100 goal blocks contain a quantifier. A quantifier is a template rather than a predicate. picking_up_trash reads:

(:goal (and
  (forall (?can__of__soda.n.01 - can__of__soda.n.01)
    (inside ?can__of__soda.n.01 ?ashcan.n.01_1))))

That is every can of soda is inside the ashcan. The parse counts one. The :objects block in the same file declares three cans. At runtime goal_status therefore holds three ground predicates, and binning two of them scores 0.67. That is partial credit which the written block does not show. This is ordinary BDDL semantics rather than a defect, which is why it appears here and not in the findings.

Two consequences for the numbers on this page. Every predicate count marked + is a floor. The denominator can also differ between instances of the same task, because each of the 40 test instances uses a different scene layout. One instance may hold three cans and another five, scored out of three and out of five.

The fix, if the numbers must be exact. Load each task instance in the simulator and read goal_status. That is about 100 scene loads at 150 to 300 s each, so 4 to 8 hours on one GPU. It would replace every + with a measured number, and allow the predicate correlation to be computed on correct counts.
Reference

All 100 tasks

Sorted shortest first. Every column states what it holds and where it comes from. Read preds against Q 2025. The first gives how finely a task can pay out. The second gives whether anyone collected.

Taskid used by --task-nameSetadded 2026, or carriedSecmean demo, length÷30 HzPredsQ denominator; + expands at runtimeBase mmetres the base droveEEF mmetres both hands movedQ 2025field mean; blank = never runSpreadlongest demo ÷ shortestDemoorganizers’ recording
turning_on_radio721670.4484.5watch ›
hanging_pictures801+760.0578.4watch ›
vacuuming_floors80156—3.8watch ›
installing_a_modem804+38—3.2watch ›
installing_smoke_detectors861910—2.9watch ›
scrubbing_bathroom_floor1051516—5.3watch ›
make_microwave_popcorn10829120.2523.1watch ›
make_cabinet_doors1151410—3.4watch ›
composting_waste1212615—2.0watch ›
clean_a_keyboard1291413—6.3watch ›
attach_a_camera_to_a_tripod130111140.0483.8watch ›
installing_a_fax_machine1362+1013—4.2watch ›
installing_a_scanner1422615—6.8watch ›
sweeping_garage1492199—2.7watch ›
make_rose_centerpieces1514920—2.6watch ›
cook_broccolini1545+1112—3.0watch ›
picking_up_trash1761+16200.3547.1watch ›
clean_a_trumpet177112190.0002.5watch ›
organizing_art_supplies20651225—2.5watch ›
set_up_a_coffee_station_in_your_kitchen209617200.0443.8watch ›
halve_an_egg21351432—2.6watch ›
chop_an_onion213417170.0393.2watch ›
spraying_for_bugs2161+22190.0652.9watch ›
boxing_food_after_dinner2235+1629—2.1watch ›
store_honey22611223—3.5watch ›
cook_a_brisket24431428—2.6watch ›
hiding_Easter_eggs2543+22230.1314.9watch ›
cook_bacon2562+21200.1201.9watch ›
store_batteries2571+1429—2.4watch ›
putting_shoes_on_rack2586+27190.2993.4watch ›
sorting_books_on_shelf25810+1024—2.5watch ›
thawing_frozen_food2749+1830—3.0watch ›
clean_boxing_gloves2741+29220.0203.5watch ›
preparing_lunch_box2755+21310.0513.2watch ›
spraying_fruit_trees278224290.0592.0watch ›
wash_a_baseball_cap2781+29220.0831.9watch ›
cook_a_frozen_pie28921934—2.5watch ›
clean_up_broken_glass2901+2928—3.1watch ›
packing_meal_for_delivery2952+2038—1.9watch ›
rearranging_kitchen_furniture2984+26310.0643.8watch ›
setting_the_fire3045+32240.0432.7watch ›
cook_hot_dogs3051+26310.1413.4watch ›
dispose_of_glass3061+1832—2.4watch ›
bringing_water3152+26200.2062.6watch ›
store_produce3172+2642—2.4watch ›
re_shelving_library_books3181+2036—2.9watch ›
make_gift_bags_for_baby_showers3223+1740—2.7watch ›
turning_out_all_lights_before_sleep3342+4034—1.9watch ›
setting_mousetraps3403+32260.2754.3watch ›
collecting_aluminum_cans3401+1842—2.6watch ›
sorting_bottles_cans_and_paper34212+2746—2.9watch ›
outfit_a_basic_toolbox355732380.0233.3watch ›
polishing_shoes3586+1741—2.2watch ›
chopping_wood358838350.0903.3watch ›
putting_dishes_away_after_cleaning3652+33600.0492.1watch ›
tidying_bedroom3683+33250.1495.4watch ›
carrying_out_garden_furniture3692+6125—1.6watch ›
bringing_paper_to_recycling37433549—2.0watch ›
wash_dog_toys3744+47320.0002.9watch ›
put_together_a_basic_pruning_kit37641937—2.4watch ›
putting_away_toys3771+3557—2.5watch ›
unloading_the_car3772+3439—2.3watch ›
can_meat3954+32540.0012.1watch ›
sorting_vegetables3975+34690.0563.1watch ›
clean_a_patio402137230.0002.6watch ›
putting_dirty_dishes_in_sink4052+5744—1.7watch ›
freeze_pies4154+35580.0143.1watch ›
cleaning_up_branches_and_twigs4182+4150—2.2watch ›
tidying_living_room4194+2137—2.6watch ›
freeze_fruit4214+2567—2.2watch ›
rearrange_your_room4282+2437—2.1watch ›
tidying_bathroom43341940—2.4watch ›
clearing_food_from_table_into_fridge4364+32330.0131.9watch ›
setup_a_bar_for_a_cocktail_party44712+2350—2.0watch ›
bringing_in_wood4511+36320.1102.1watch ›
cleaning_up_plates_and_food4564+36340.0942.4watch ›
putting_up_Christmas_decorations_inside4577+41390.0862.0watch ›
putting_away_Halloween_decorations4604+47400.148127.7watch ›
cook_cabbage472439440.0002.6watch ›
carrying_in_groceries4764+37400.0501.8watch ›
laying_tile_floors4785+6544—1.7watch ›
dispose_of_batteries4812+4733—3.4watch ›
organizing_school_stuff48662652—2.8watch ›
moving_boxes_to_storage486437190.3832.1watch ›
slicing_vegetables4957+41400.0302.9watch ›
clean_your_rusty_garden_tools50556948—3.0watch ›
cook_brussels_sprouts5225+3563—2.4watch ›
getting_organized_for_work5221047440.0032.9watch ›
sorting_household_items5277+51540.0192.8watch ›
stacking_wood5316+8643—3.3watch ›
setting_the_table5944+5685—2.0watch ›
picking_up_toys6303+47420.0414.2watch ›
collecting_childrens_toys6404+58500.1282.9watch ›
make_pizza640257780.0002.6watch ›
loading_the_car641448410.0362.4watch ›
storing_food6624+56580.0521.6watch ›
clean_up_your_desk7148+60740.0152.3watch ›
canning_food7669+64920.0012.5watch ›
boxing_books_up_for_storage8081+72500.0112.4watch ›
assembling_gift_baskets8694+80790.0512.4watch ›