BFP is a goal-conditioned visuomotor imitation policy that extrapolates to unseen goals while preserving multimodal behaviour. The key idea: retrieve a training example (the anchor) such that the anchor and the residual to the current observation make action prediction simple, and generate actions with our bilinear flow. The key result: the whole action distribution extrapolates, the theory yields pre-rollout diagnostics of which checkpoints and goals will do well, and BFP generalizes to unseen goals on a real robot.
| Success on unseen goals | Simulation | Real robot |
|---|---|---|
| Flow[2] | 30% | 18% |
| BRP[1] | 58% | 39% |
| BFP | 79% | 50% |
Goal-conditioned imitation learning (GCIL) with flow matching is a promising framework that can represent multimodal behaviors while adapting to diverse, user-specified goals, yet often fails when goals lie outside the demonstration support. To extrapolate to such unseen goals without collapsing multimodality—a problem we call distributional extrapolation—we introduce Bilinear Flow Policy (BFP), a generative visuomotor policy that combines transductive retrieval with a bilinear conditional flow.
Given an unseen observation-goal pair, BFP retrieves an “anchor” training example and transductively reformulates the unseen pair as this familiar anchor plus a residual term. For this decomposition to guide action prediction, the residual must compactly encode how the current observation-goal pair differs from the anchor, and the anchor must be chosen so that this difference is predictive of the corresponding action distribution. BFP achieves this with pretrained visual features and a novel learned anchor-selection algorithm. The novel bilinear flow then models how the anchor and the residual jointly determine the multimodal action distribution. We prove that, for bilinear flow, action distribution error at unseen goals is bounded by the in-distribution flow-matching error up to problem-dependent factors.
Across five manipulation tasks in simulation, BFP achieves 2.63× the out-of-distribution success rate of a GCIL policy and 1.36× that of the strongest extrapolation-targeted baseline. On two real-world tasks, BFP improves over GCIL by 32 percentage points on average. Finally, our theory yields practical, pre-deployment diagnostics for predicting which trained policies will extrapolate well and to which unseen goals.
Every unseen-goal trial behind the paper’s hardware results, uncurated. Pick a goal, a plate position or a rack hole, and watch BFP and Flow attempt it.
Platform. A 7-DoF arm with a parallel-jaw gripper, a fixed RGB-D scene camera and an RGB-D wrist camera, controlled at 20 Hz. Depth is used only to extract the goal, never by the policy. Demonstrations are teleoperated with a VR headset. The goal is extracted from the scene camera and expressed in the robot base frame: an open-vocabulary segmenter finds the target object and a 6-DoF pose tracker registers it.
PnPMug. Pick up a black mug and place it on a green plate; the goal is the plate’s 3-D position. 127 demonstrations with the plate restricted to one region; for evaluation the plate is placed only in the held-out region, so no goal lies within the bounding box of the demonstrated goals. 50 unseen initial conditions with 2 trials each per policy, plus 50 in-distribution conditions with 1 trial each; episodes time out at 30 s.
InsertPipette. Pick a pipette from a cradle and insert it into the test tube in the goal hole of a 4 × 8 rack with 36 mm pitch; the goal is the hole’s discrete coordinate. 133 demonstrations span a 2 × 5 block of training holes. The remaining 22 holes are evaluated with 5 trials each per policy, the training holes with 3; episodes time out at 60 s.
Policies. All policies share the same frozen pretrained visual encoder, both camera views, proprioception and the explicit goal, and output absolute end-effector poses and gripper commands as 16-step action chunks from a 1-D U-Net backbone, executing 8 steps per chunk with temporal ensembling. Flow is the same policy without the anchor–residual decomposition and the bilinear interaction. One trained seed per method. A human evaluator scores binary success; a normalized dense reward tracks task progress (reach, grasp, lift, place, retreat for the mug; reach, grasp-and-lift, align, insert, retreat for the pipette).
Flow policies imitate multimodal behaviour well, but only among the goals they were shown. We want the same behaviour at goals nobody demonstrated.
Goal-conditioned imitation learning with flow matching[2] is powerful. One policy can represent several valid ways of doing a task and pick among them according to a user-specified goal. It is also brittle: it imitates what it was shown, and only where it was shown.
At deployment, the goal is often one that no demonstration covered. For example, the plate is placed further to the side than in any demonstration; the test tube sits in a hole of the rack that was never used. Once the observation–goal pair leaves the demonstrated support, a flow policy has nothing to imitate, and its behaviour degrades in ways that are hard to predict.
How can a policy generalize to an unseen goal while preserving the multimodality of the expert? We call this the distributional extrapolation problem. Each half is easy in isolation, and neither is enough on its own. Regressing a single action can extrapolate, but it averages the expert’s modes into an action nobody demonstrated. Modelling the full action distribution keeps the modes, but says nothing about how they should continue beyond the data.
Let \(\mathcal{A}\subset\mathbb{R}^{d_a}\) be the action space and \(\mathcal{A}^{H}\) the space of \(H\)-step action chunks. We are given expert demonstrations \(\mathcal{D}=\{\tau^{(i)}\}_{i=1}^{N}\), \(\tau^{(i)}=\{(\mathbf{x}^{(i)}_t,\mathbf{a}^{(i)}_t)\}_{t=1}^{T_i}\), where the policy condition \(\mathbf{x}_t=(\mathbf{o}_t,\mathbf{g})\) pairs a state or visuomotor observation with a compact, explicitly specified episode goal such as a target object pose. Each timestep defines an action-chunk target \(A_t=(\mathbf{a}_t,\dots,\mathbf{a}_{t+H-1})\).
Let \(p_{\mathrm{tr}}\) and \(p_{\mathrm{te}}\) be the marginal distributions of conditions under training and evaluation. Conditions inside the demonstrated support are in-distribution (ID); conditions outside it are out-of-distribution (OOD). The expert conditional \(q^\star(\cdot\mid\mathbf{x})\) is the same in both regimes, but \(p_{\mathrm{te}}\) may contain conditions induced by unseen goals outside \(\operatorname{supp}(p_{\mathrm{tr}})\). Given only \(\mathcal{D}\), we seek a policy \(\pi\) with small test distributional risk
where \(W_2^2\) is the squared 2-Wasserstein distance. A policy learned from demonstrated conditions must match the full expert action distribution at unseen conditions. This matters exactly when a condition admits several valid action chunks: mean regression averages the modes, and distributional modelling alone does not specify how the modes should extend beyond the training support.
Retrieve a training example (the anchor), describe the unseen condition as that anchor plus a residual, and let a bilinear flow turn the pair into a full action distribution.
The idea comes from transduction. Faced with a query it has never seen, the policy retrieves an anchor from the demonstrations, describes the query by its residual to that anchor, and predicts from the pair. Bilinear transduction (BT)[1] did this for deterministic regression in low-dimensional state spaces with hand-designed anchor rules.
BFP carries the recipe to visuomotor imitation with three parts: a pretrained representation in which residuals are meaningful, a learned anchor policy that picks anchors that make prediction easy, and a bilinear conditional flow that generates the full multimodal action distribution.
A rank-\(r\) bilinear head predicts the flow’s velocity field from the anchor and the residual.
For a query \(\mathbf{x}=(\mathbf{o},\mathbf{g})\) and a retrieved anchor \(\mathbf{x}^\dagger\), both conditions are encoded with the same representation \(\phi\), giving \(\mathbf{z}=\phi(\mathbf{x})\), \(\mathbf{z}^\dagger=\phi(\mathbf{x}^\dagger)\) and the residual \(\boldsymbol{\delta}=\mathbf{z}-\mathbf{z}^\dagger\). BFP models the flow-matching velocity of each action coordinate \(k\) as a rank-\(r\) inner product of a residual branch and an anchor branch:
Both branches may be nonlinear; only their interaction is low-rank. Because the structure is imposed on the velocity field, the generative model inherits the extrapolation of bilinear transduction and keeps every mode of the expert. The head is trained with the ordinary conditional flow matching loss on anchor–residual–action triples.
Frozen pretrained features give query and anchor a common coordinate system; a learned scorer picks the anchors the bilinear head can explain.
Images are encoded by a frozen pretrained visual encoder (R3M[3] or DINOv3[4]) and concatenated with proprioception and the goal. Because the encoder is trained independently of the demonstrations, residuals live in a coordinate system shared by train and test, where they track changes of the condition and are insensitive to nuisance appearance.
For each query, a coarse phase estimator proposes a window of candidate anchors from demonstrations at a similar task progress, and a learned scorer places a softmax over them. Predictor and scorer are trained in alternation: the predictor learns mostly from the anchors the scorer favours, and the scorer learns to favour anchors with low held-out BFP loss, that is, anchor–residual pairs the shared rank-\(r\) head explains well.
At deployment, choose anchors whose residual was covered during training.
At test time the aim changes from making prediction easy to keeping the residual on support. Following BT, BFP first keeps the demonstrations whose goal residual to the query is closest to a goal residual seen in training, then applies the learned scorer to those candidates and samples an anchor. With the anchor fixed, it integrates the bilinear flow from noise to an action chunk.
The guarantee rests on two requirements, inherited from bilinear transduction and restated for the velocity field. (S) Bilinear structure: the anchor is selected, and the residual defined, so that the anchor–residual-to-velocity problem is approximately low-rank. (C) Residual coverage: the test-time anchor yields a residual that was covered during training. Write \(\mathcal{E}^{\mathrm{flow}}_{s}(\theta)=\mathbb{E}_{s}\big[\|v_\theta-u^\star\|_2^2\big]\) for the flow error on the training (\(s=\mathrm{tr}\)) or test (\(s=\mathrm{te}\)) conditions.
Theorem (distributional extrapolation of BFP, informal). Suppose the marginal velocity field satisfies (S) and test-time anchor selection satisfies (C). Under the remaining technical conditions detailed in the paper’s appendix, sufficiently small training flow error implies
Reading it. In-distribution flow error controls the full generated action distribution at unseen goals, provided the representation and retrieval expose a low-rank and covered factorization. The first inequality converts velocity-field error into Wasserstein error of the generated distribution; the second is the bilinear transfer from training to test conditions. Both requirements are measurable, which turns the theorem into the two pre-rollout diagnostics used below: held-out flow error for (S) and residual coverage for (C).
BFP extrapolates both expert modes on toy problems, has the highest unseen-goal success on every simulated setting, and its in-distribution diagnostics predict which checkpoints and which goals will extrapolate.
With representation and retrieval held fixed, only the bilinear flow carries both expert modes past the training support.
Three one-dimensional bimodal problems isolate the predictor from everything else: inputs are already compact and random anchors suffice. Flow matches the expert in distribution, but its modes drift once the input leaves the training range. Deterministic bilinear regression (BRP) extrapolates structurally but collapses both modes onto their average. BFP tracks both modes and keeps their balance beyond the support. The table carries the numbers; an oracle that samples the expert directly sets the floor.
| Method | NW2 ID ↓ | NW2 OOD ↓ | Balance OOD ↓ |
|---|---|---|---|
| Oracle | 0.078 | 0.086 | 0.019 |
| Flow | 0.221 | 0.506 | 0.160 |
| BRP | 0.726 | 0.763 | 0.500 |
| BFP | 0.224 | 0.321 | 0.103 |
Across seven visuomotor settings BFP has the highest out-of-distribution success on every one and forms the top statistical group, without giving up in-distribution success.
Five manipulation tasks give seven settings (two tasks also come with a larger set of initial conditions), each evaluated on fixed sets of in- and out-of-distribution goals over three training seeds, with family-wise error control for the comparisons. We compare with Flow, the standard generative GCIL policy, and four extrapolation-targeted baselines: EquiDP[5] prescribes a geometric equivariance; DARP[6] predicts from the nearest anchors with a local, non-bilinear model; BRP[1] uses time-synchronized anchors and deterministic bilinear regression; CTA[7] learns a task-endogenous analogy and uses it as the residual.





The comparisons say where the gain comes from. Learning the extrapolation relation from demonstrations beats prescribing it, even against EquiDP’s carefully tuned, task-specific rule. Local retrieval alone (DARP) or a learned analogy alone (CTA) does not match the structured predictor. And the margin over BRP is what learned retrieval and generative bilinear flow add on top of the fixed-anchor, deterministic recipe. The same ordering holds with state observations.
Each panel shows the support of the training conditions (left) and of the held-out evaluation conditions (right) from the same camera, so the shift is directly visible.





For BFP, yes: held-out in-distribution flow error ranks checkpoints by their out-of-distribution error, and residual coverage ranks queries. For plain Flow, in-distribution fit says nothing about OOD.
The guarantee has two measurable ingredients, so we use them as diagnostics before any out-of-distribution rollout. First, rank checkpoints by held-out in-distribution flow error. For Flow, a checkpoint that fits held-out demonstrations better is not systematically better out of distribution. For BFP the ranking carries over. Second, within a fixed BFP checkpoint, group queries by how well their retrieved residual is covered by the training residuals: poorly covered queries incur markedly higher flow error than well-covered ones.
The practical upshot is direct. Select BFP checkpoints and hyperparameters on in-distribution validation loss, and prefer anchors that keep the residual on support. Probe construction and the per-task numbers are in the paper.
Frozen pretrained features and learned phase-conditioned retrieval matter most. A stronger backbone helps BFP and hurts Flow. Rank barely matters.
A frozen pretrained encoder beats training the visual encoder jointly from scratch by a wide margin, and learned phase-conditioned retrieval beats both time-synchronized retrieval (as in BT) and random retrieval. One might worry that BFP extrapolates only because the low-rank head underfits, and that a more expressive architecture would erase the gain by overfitting. The backbone sweep says otherwise: swapping the MLP for a sequence-aware Transformer improves BFP and degrades Flow, with near-perfect in-distribution success across variants. Out-of-distribution behaviour is governed by the bilinear interaction through which the backbone combines anchor and residual, and a stronger backbone sharpens it. Performance is insensitive to the rank.
BFP assumes a compact, explicitly specified goal: a plate position, a hole index. That assumption is what makes the residual meaningful and the test-time coverage rule simple. Relaxing it is the next step.