Existing evaluations of generative inverse solvers miss critical failures like mode collapse and overconfident uncertainty—PosteriorBench reveals these gaps by directly comparing predicted solution distributions to ground-truth posteriors.
PosteriorBench is a benchmark for evaluating how well generative models solve inverse problems by checking if they capture the full range of possible solutions, not just single best guesses. It tests four physics problems using reference posteriors from MCMC and rejection sampling, with metrics measuring accuracy, uncertainty, and distributional fit.