Real-world evaluation is also difficult to standardize, often requiring human resets and suffering from environment variability. As a result, simulation-based benchmarks have become popular for their reproducibility and ease of use.
Existing benchmarks mostly focus on household tasks. However, retail and logistics scenarios — such as shelf picking or order packing — remain underexplored. Dedicated benchmarks for these domains are needed to advance robotic capabilities in retail environments.
RoboBenchMart addresses limitations of prior works by providing code to generate diverse store layouts and robotic trajectories, enabling the training and benchmarking of robotic policies in retail environments.
| Benchmark/ Dataset |
Published in |
Retail Domain |
Scene Generation |
Arrangement Generation |
Release 3D Assets |
Trajectories Generation |
Tasks Diversity |
Atomic Tasks |
Composite Tasks |
|---|---|---|---|---|---|---|---|---|---|
| ALFRED | CVPR'20 | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| RLBench | RA-L'19 | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| RoboCasa | RSS'24 | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| CALVIN | RA-L'22 | ✗ | ✗ | ✓ | n/a | ✓ | ✓ | ✓ | ✓ |
| LIBERO | NeurIPS'23 | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| VLABench | arXiv'24 | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| BEHAVIOR-1K | CoRL'22 | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| ManiSkill-HUB | ICLR'25 | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| RP2K | arXiv'20 | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| SKU110K | CVPR'19 | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| StandardSim | ICIAP'22 | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| IPA-3D1K | IROS'23 | ✓ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ |
| FetchBot | arXiv'25 | ✓ | ✗ | ✓ | ✗ | ✓ | ✗ | ✓ | ✗ |
| RoboBenchMart | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Model | Param. (#) | Testing scenario |
Atomic Tasks | Composite Tasks | |||||
|---|---|---|---|---|---|---|---|---|---|
| Pick to basket |
Pick from floor |
From board to board |
Open fridge |
Close fridge |
Pick 3 items |
Pick from fridge |
|||
| Octo | 93M | In-Domain | 21 | 1 | 21 | 27 | 45 | 0 | 0 |
| Unseen Scenes | 2 | 1 | 1 | 23 | 26 | 0 | 0 | ||
| Unseen Scenes & Items | 0 | 0 | 0 | n/a | n/a | 0 | 0 | ||
| SmolVLA | 450M | In-Domain | 0 | 0 | 0 | 10 | 13 | 0 | 0 |
| Unseen Scenes | 0 | 0 | 0 | 13 | 13 | 0 | 0 | ||
| Unseen Scenes & Items | 0 | 0 | 0 | n/a | n/a | 0 | 0 | ||
| π0 | 3.3B | In-Domain | 14 | 21 | 19 | 61 | 95 | 0 | 0 |
| Unseen Scenes | 1 | 6 | 3 | 44 | 85 | 0 | 0 | ||
| Unseen Scenes & Items | 0 | 1 | 0 | n/a | n/a | 0 | 0 | ||
| π0.5 | 3.3B | In-Domain | 55 | 21 | 56 | 56 | 91 | 0 | 0 |
| Unseen Scenes | 31 | 4 | 27 | 49 | 78 | 0 | 0 | ||
| Unseen Scenes & Items | 5 | 3 | 21 | n/a | n/a | 0 | 0 | ||
The success-rate table shows that current generalist VLA models struggle even with basic retail tasks. SmolVLA performs poorly across all scenarios and succeeds only on the simplest fridge-opening and fridge-closing tasks, which may reflect its more limited pretraining data. Octo and $\pi_{0}$ reach moderate performance in the In-Domain setting, but their success rates degrade sharply on Unseen Scenes and collapse in Unseen Scenes & Items. $\pi_{0.5}$ performs best and is the only model with non-zero success in Unseen Scenes & Items, although its performance is still far from reliable. All models achieve zero success on composite tasks, suggesting limited ability to execute multi-step instructions and generalize across task stages.
| task | target | move | pregrasp | grasp | drop | displace | preplace | place-coord | place | partial | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Octo | 7 | 22 | 8 | 9 | 39 | 2 | 2 | 1 | 1 | 6 | 3 |
| SmolVLA | 4 | 29 | 16 | 32 | 17 | 0 | 0 | 0 | 0 | 0 | 2 |
| π0 | 0 | 23 | 1 | 9 | 43 | 5 | 1 | 2 | 4 | 5 | 7 |
| π0.5 | 0 | 20 | 1 | 2 | 50 | 4 | 7 | 0 | 2 | 9 | 5 |
The failure-rate table highlights two main bottlenecks. Octo, $\pi_{0}$, and $\pi_{0.5}$ fail predominantly at the grasp stage, indicating that precise shelf-level interaction remains difficult even when the correct region is reached. SmolVLA most often fails earlier, during pregrasp positioning. Also SmolVLA and Octo shows movement-related failures when approaching target fixtures. Target grounding is another common issue across models, showing that selecting the instructed product in cluttered shelves is still unreliable.
Overall, these results highlight three limitations of current generalist models: fragility to minor scene changes such as layouts, textures, and object placements; poor generalization from limited demonstrations to novel object-task combinations; and insufficient reliability for shelf-level grasping and compositional execution. This suggests that existing pretrained models may be insufficient for effective adaptation in retail environments without more targeted retail-specific pretraining.
@article{soshin2025robobenchmart,
title={RoboBenchMart: Benchmarking Robots in Retail Environment},
author={Soshin, Konstantin and Krapukhin, Alexander and Spiridonov, Andrei and Bukhtuev, Gregorii and Kuznetsov, Andrey and Shakhuro, Vlad and Shepelev, Denis},
journal={arXiv preprint arXiv:2511.10276},
year={2025}
}