MUNITEUnified Multimodal Latent Inference for Any-to-Any Multimodal Generation

Kyeongmin Yeo Minhyuk Sung

KAIST

{aaaaa, mhsung}@kaist.ac.kr

Code coming soon

Skip Figure 1
One model, three observation regimes A single conditional flow transports Gaussian noise to the latent marginal with no observations, a conditional distribution with partial observations, or one encoded point with full observations. An image, audio and text are decoded from one shared latent sample. The densities are schematic, and the displayed outputs form one unconditional sample. ConditioningLatent spaceGeneration Unconditional example Unconditional Partial observation Full observation Image Audio Text “The sun sets overa town onthe coast.” One model, three observation regimes A single conditional flow transports Gaussian noise to the latent marginal with no observations, a conditional distribution with partial observations, or one encoded point with full observations. An image, audio and text are decoded from one shared latent sample. The densities are schematic, and the displayed outputs form one unconditional sample. Conditioning Latent space Unconditional Partial Full Generation Image Audio Text “The sunsets overa town onthe coast.” Unconditional example

Latent inference and decoding

Marginal sampling, conditional inference and full-observation encoding use the same latent space and model.

Figure 1. Schematic latent distributions. The displayed outputs form one unconditional sample.

Abstract

We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently.

To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising.

Across PolyMNIST-D-Q, FFHQ64, and image–text–audio, MUNITE achieves competitive or better generation quality and source–target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image–text–audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.

Method

Generating target modalities

Given the observed modalities, MUNITE samples one latent w for all target decoders. The outputs are generated independently given w, so the latent must capture cross-modal dependence. Each decoder models the remaining uncertainty in its modality.

MUNI samples a shared latent from a diagonal-Gaussian posterior aggregated from the observed modalities. MUNITE samples from the conditional distribution of the full-observation latent with a flow trained by conditional flow matching.

Factorization
Pdata(dxT∣xS)=∫ [∏m∈TPdata(dxm∣W=w)] QE(dw∣xS)

𝒮 and 𝒯 index disjoint observed and target modalities. Q𝓔(· | x𝒮) is the conditional distribution of the full-observation latent W given x𝒮.

This factorization assumes conditional independence given W.

Training objectives

The shared inference network and modality-specific decoders are trained jointly on complete and incomplete examples.

Skip Figure 2

Reconstruction with target detachingOther observed inputs and a Gaussian query feed one shared inference network and a modality-specific decoder. The observed target also supplies the loss. Its conditional input edge, when m is in S, keeps forward values but stops reconstruction gradients through target keys and values.

Decoders reconstruct observed modalities from the predicted latent. Stopping reconstruction gradients through a target's own input discourages copying modality-private information.

Self-distillation at a common detached stateFor incomplete examples, an N-step flow conditioned on all available modalities constructs the noisy latent state. At that detached state and time, the same network predicts clean latents from the available modalities and a smaller subset. The teacher prediction is detached before the loss.TeacherStudent

Predictions conditioned on more modalities supervise the same network conditioned on fewer. This trains conditional latent inference from incomplete examples.

Symmetric contrastive alignment of complementary subsetsStacked observations and model boxes denote batch examples. Nonempty complementary S and C use independent Gaussian queries through one shared network. Flattened clean predictions are l2-normalized. Symmetric InfoNCE uses same-example positives and other batch examples as negatives.

A contrastive loss aligns latent predictions from complementary subsets of the same example, using other batch examples as negatives.

Figure 2. Three objectives jointly train the shared inference network.
Proposition 1: gradient equivalence

An exact teacher predicts μ𝒜, the conditional mean of the full-observation latent given the noisy state and the available observations. Replacing W with this prediction changes the expected loss by a term independent of θ. The expected gradient through the student's prediction is unchanged.

W is the full-observation latent and Wt its noisy interpolant. The teacher observes 𝒜, and the student observes a subset 𝒮 ⊊ 𝒜. The two unweighted losses compare the student's clean-latent prediction with W and μ𝒜, respectively. θ denotes the student parameters.

The representation and teacher are held fixed. We assume that the teacher's flow reproduces the true conditional marginals of Wt and that unavailable modalities are removed independently of the data. In practice, the current network serves as the teacher, and its flow is integrated numerically.

Results

MUNITE achieves the highest coherence in all six joint-generation comparisons.

Coherence between jointly generated modalities. ↑ Higher is better. Best values are bold.
MethodOne-to-manyUnconditional
Text → I, AImage → T, AAudio → T, IT, IT, AI, A
AIS ↑CLAP ↑CLIP ↑CLIP ↑CLAP ↑AIS ↑
CoDi63.8667.41823.414———
OmniFlow77.02714.78122.69721.1714.2350.95
MUNI81.55832.23725.85026.94928.46781.989
CFM74.70221.69524.03423.62219.89674.208
DFM73.63520.95124.03723.74019.50371.942
FlowBind79.81925.81224.853———
MUNITE83.33634.76226.33427.39830.22087.045

CLIP: text–image · CLAP: text–audio · AIS: image–audio. Scores ×100. A dash (—) marks an unsupported setting. OmniFlow unconditional scores are from MUNI (Yeo et al., 2026).

Qualitative results

PolyMNIST-D-Q

Coherence measures pairwise agreement on the unobserved digit or quadrant.

Observed: Bottom-left quadrant

Unobserved: digit

View 0View 1View 2MUNIMUNI view 0, conditioned on Bottom-left quadrantMUNI view 1, conditioned on Bottom-left quadrantMUNI view 2, conditioned on Bottom-left quadrantMUNITEMUNITE view 0, conditioned on Bottom-left quadrantMUNITE view 1, conditioned on Bottom-left quadrantMUNITE view 2, conditioned on Bottom-left quadrant
The MUNITE views shown share the unobserved digit, while the MUNI views differ.

Coherence of the unobserved digit

MUNI0.1074
MUNITE, separate draws0.1014
MUNITE0.9904

Independent-uniform chance: 0.10.

Observed: Digit 1

Unobserved: quadrant

View 0View 1View 2MUNIMUNI view 0, conditioned on Digit 1MUNI view 1, conditioned on Digit 1MUNI view 2, conditioned on Digit 1MUNITEMUNITE view 0, conditioned on Digit 1MUNITE view 1, conditioned on Digit 1MUNITE view 2, conditioned on Digit 1
The MUNITE views shown share the unobserved quadrant, while the MUNI views differ.

Coherence of the unobserved quadrant

MUNI0.2528
MUNITE, separate draws0.2500
MUNITE1.0000

Independent-uniform chance: 0.25.

Separate draws use independent latent samples for the generated views.

FFHQ64

Age and gender labels condition joint generation of RGB, semantic segmentation and surface normals.

Observed age and gender
RGBSegmentationNormalsSeg. | RGBMUNIMUNI RGB, conditioned on Male, 0–2MUNI segmentation, conditioned on Male, 0–2MUNI surface normals, conditioned on Male, 0–2
MUNITEMUNITE RGB, conditioned on Male, 0–2MUNITE segmentation, conditioned on Male, 0–2MUNITE surface normals, conditioned on Male, 0–2
Sample 1 of 8 · Male, 0–2

Image–text–audio

Unconditional generation and three one-to-many routes.

Generated image and audio features are rendered with pretrained models, following MUNI (Yeo et al., 2026).

BibTeX

@misc{yeo2026muniteunifiedmultimodallatent,
  title={MUNITE: Unified Multimodal Latent Inference for Any-to-Any Multimodal Generation},
  author={Kyeongmin Yeo and Minhyuk Sung},
  year={2026},
  eprint={2610.09866},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2610.09866},
}