Observed: Bottom-left quadrant
Unobserved: digit
Coherence of the unobserved digit
Independent-uniform chance: 0.10.
Code coming soon
Marginal sampling, conditional inference and full-observation encoding use the same latent space and model.
We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently.
To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising.
Across PolyMNIST-D-Q, FFHQ64, and image–text–audio, MUNITE achieves competitive or better generation quality and source–target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image–text–audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.
Given the observed modalities, MUNITE samples one latent w for all target decoders. The outputs are generated independently given w, so the latent must capture cross-modal dependence. Each decoder models the remaining uncertainty in its modality.
MUNI samples a shared latent from a diagonal-Gaussian posterior aggregated from the observed modalities. MUNITE samples from the conditional distribution of the full-observation latent with a flow trained by conditional flow matching.
Scroll horizontally to read the full equation.
𝒮 and 𝒯 index disjoint observed and target modalities. Q𝓔(· | x𝒮) is the conditional distribution of the full-observation latent W given x𝒮.
This factorization assumes conditional independence given W.
The shared inference network and modality-specific decoders are trained jointly on complete and incomplete examples.
Decoders reconstruct observed modalities from the predicted latent. Stopping reconstruction gradients through a target's own input discourages copying modality-private information.
Predictions conditioned on more modalities supervise the same network conditioned on fewer. This trains conditional latent inference from incomplete examples.
A contrastive loss aligns latent predictions from complementary subsets of the same example, using other batch examples as negatives.
An exact teacher predicts μ𝒜, the conditional mean of the full-observation latent given the noisy state and the available observations. Replacing W with this prediction changes the expected loss by a term independent of θ. The expected gradient through the student's prediction is unchanged.
W is the full-observation latent and Wt its noisy interpolant. The teacher observes 𝒜, and the student observes a subset 𝒮 ⊊ 𝒜. The two unweighted losses compare the student's clean-latent prediction with W and μ𝒜, respectively. θ denotes the student parameters.
The representation and teacher are held fixed. We assume that the teacher's flow reproduces the true conditional marginals of Wt and that unavailable modalities are removed independently of the data. In practice, the current network serves as the teacher, and its flow is integrated numerically.
MUNITE achieves the highest coherence in all six joint-generation comparisons.
| Method | One-to-many | Unconditional | ||||
|---|---|---|---|---|---|---|
| Text → I, A | Image → T, A | Audio → T, I | T, I | T, A | I, A | |
| AIS ↑ | CLAP ↑ | CLIP ↑ | CLIP ↑ | CLAP ↑ | AIS ↑ | |
| CoDi | 63.866 | 7.418 | 23.414 | — | — | — |
| OmniFlow | 77.027 | 14.781 | 22.697 | 21.17 | 14.23 | 50.95 |
| MUNI | 81.558 | 32.237 | 25.850 | 26.949 | 28.467 | 81.989 |
| CFM | 74.702 | 21.695 | 24.034 | 23.622 | 19.896 | 74.208 |
| DFM | 73.635 | 20.951 | 24.037 | 23.740 | 19.503 | 71.942 |
| FlowBind | 79.819 | 25.812 | 24.853 | — | — | — |
| MUNITE | 83.336 | 34.762 | 26.334 | 27.398 | 30.220 | 87.045 |
CLIP: text–image · CLAP: text–audio · AIS: image–audio. Scores ×100. A dash (—) marks an unsupported setting. OmniFlow unconditional scores are from MUNI (Yeo et al., 2026).
Coherence measures pairwise agreement on the unobserved digit or quadrant.
Unobserved: digit
Coherence of the unobserved digit
Independent-uniform chance: 0.10.
Unobserved: quadrant
Coherence of the unobserved quadrant
Independent-uniform chance: 0.25.
Separate draws use independent latent samples for the generated views.
Age and gender labels condition joint generation of RGB, semantic segmentation and surface normals.
Unconditional generation and three one-to-many routes.
Generated image and audio features are rendered with pretrained models, following MUNI (Yeo et al., 2026).
No modality is observed. An image, a caption and a sound are generated together.

“The sun sets over a town on the coast.”
A13 · Sample 1 of 5

“a wolf howling with light rustling in the background”
A14 · Sample 2 of 5

“Two dogs that are looking at each other.”
A15 · Sample 3 of 5
Text is observed. An image and a sound are generated together.
“vibrations from a sewing machine”

A23 · Sample 1 of 5
“a fire siren with people talking”

A24 · Sample 2 of 5
“a woman speaks followed by birds chirping”

A25 · Sample 3 of 5
An image is observed. A caption and a sound are generated together. The observed images were generated with FLUX.1.

“a bell rings continuously”
A28 · Sample 1 of 5

“A horse is standing in the grass with lightning behind it.”
A29 · Sample 2 of 5

“a clock tick-tocks”
A30 · Sample 3 of 5
A sound is observed. A caption and an image are generated together.

“church bells ringing in the distance”
A33 · Sample 1 of 5

“a group of ducks quacking as a crowd of people talk in the background”
A34 · Sample 2 of 5

“a man gives a speech then an audience gives applause”
A35 · Sample 3 of 5
@misc{yeo2026muniteunifiedmultimodallatent,
title={MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation},
author={Kyeongmin Yeo and Minhyuk Sung},
year={2026},
eprint={2610.09866},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2610.09866},
}