Overview
Low-confidence remasking can improve individual answers from diffusion language models while narrowing the range of reasoning paths they explore. This work explains that trade-off and develops an Independent Metropolis–Hastings sampler to balance generation quality with diversity. Experiments evaluate mathematical reasoning and code generation across MATH500, AIME24/25, HumanEval, and MBPP.