A masked-diffusion language model (iLLaDA-8B) fine-tuned with LoRA to solve 9×9 Sudoku, compared against
autoregressive LLMs. Diffusion fills the whole grid in parallel and unmasks the cells it is most confident about first,
so it can use constraints from anywhere on the board instead of committing left to right.
Adapter: Techno03/illada-8b-sudoku-lora ·
Code: github.com/arush-garg/SudokuDiffusion
| Difficulty | iLLaDA zero-shot | Gemma-4 zero-shot | iLLaDA fine-tuned | Gemma-4 fine-tuned | Llama 3.3 70B |
|---|---|---|---|---|---|
| easy (n=13) | ~35.8% | 72.5% | 92.7% | 95.7% | 35.2% |
| medium (n=70) | ~8.2% | 21.5% | 69.6% | 38.9% | 14.4% |
| hard (n=17) | ~0.0% | 3.6% | 27.4% | 8.9% | 7.5% |
These are the real outputs recorded during evaluation for the 100 held-out puzzles, scored cell by cell. Live inference isn't hosted here because the 8B model needs a GPU.