Each iteration produces some visual result. Which must feed the next iteration loop. I have tried 3 trivial scenarios - a) rasterize HTML and transduce into text using Visual AI model b) AI text model reads DOM c) You share it with designer, designer converts them into text instructions.
Visual AI samples and doesn't enumerate problems, bad at global consistency/invariants
Raw DOM - text from DOM is low SNR proxy for visual state with
Designer - slow loop
reply