
Instructions are never complete specifications of the behavior they request: speakers say only what their listeners cannot supply for themselves (Grice, 1975; Zipf, 2016). An agent told to “prepare an apple” must decide unaided whether to wash or cut it and where to place it. Such choices differ between people and are seldom stated, only enacted (Slovic, 1995; Lichtenstein and Slovic, 2006). Multimodal foundation models render instructions into grounded plans (Brohan et al., 2023; Driess et al., 2023), yet settle these residual choices by generic priors rather than by evidence about the person served (Zhang et al., 2024). Whether an agent can instead recover such constraints from prior behavior has not been separated from the easier question of planning once they are supplied. Here we show the failure to be largely one of acquisition rather than planning, and that stating the inferred preference before acting recovers much of what is lost. Preference-based Planning (PBP) benchmarks 290 preferences over action parameters, placement policies, and temporal orderings, asking agents to infer from a few unlabeled demonstrations which constraint governs a new underspecified instruction. Given the preference, the strongest models plan categorical tasks at near-zero error; required to infer it, performance drops sharply, and without demonstrations accuracy falls by over a third. Because a verbalized rule keeps intent while discarding perceptual particulars, preferences so represented transfer to unfamiliar scenes, whereas implicit imitation stays bound to the rooms it was learned in. The bottleneck lies not in generating actions but in the pragmatic inference preceding them, for which language proves an effective medium.