[R] Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation
TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity.
Authors: Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher, Claire Vernade
Arxiv: https://arxiv.org/abs/2605.12000
Github: https://github.com/ziyadsheeba/mabc
https://preview.redd.it/i20adc3z04uh1.png?width=2532&format=png&auto=…
scorecomments4 sightings
first seen 2026-10-07 21:24 UTClast seen 2026-10-08 11:29 UTCscore then 1score now 3gained +2sightings 4
open on reddit ↗
💬 0