coding models have historically done poorly at ui generation tasks, and while they have gotten much better recently, they still struggle in certain domains or use cases outside its training distribution. roblox ui is one such area, as it is written in luau against roblox-specific apis and has unique stylistic requirements.
we worked with lemonade, who is building a roblox coding harness, to post-train hy3, a 295b model, to more reliably build visually appealing roblox ui for their users needs.
context
roblox studio is the software that developers use to create games on the roblox game engine. lemonade is an agentic tool that empowers users to build their roblox games with just simple prompts, whether it be building a game from scratch, adding a new game mechanic to their project. with lemonade, a developer describes what they want and the lemonade agent interfaces with the roblox studio instance to build it in their place: its able to navigate the workspace, create instances, edit scripts, and run playtests.
there’s a wide range of ui that users asks for, and lemonade strives to maintain visual consistency across variety of prompts, even with little detail from their users. they have specific style templates that the user can specify. here are some examples of generated ui using the default studded blocky theme:




outputs by grok 4.5
lemonade relies on frontier models to power their agent, and found that most are quite bad at generating roblox ui. across their evals, they found grok 4.5 to be capable, but the problem is cost at scale. to scale to their hundreds of thousands of users, they needed a cheaper offering, but smaller models were barely capable:


outputs by a 35b model
given that, the goal became: can we post-train a smaller model to be just as good as grok 4.5 for generating roblox ui in their harness.
building the environment
rl requires sandboxed environments to run their rollouts in parallel, and we typically run up to 128. lemonade’s agent builds and tests the game inside roblox studio, but studio is a desktop application that bundles the full game engine, doesn’t run in linux, and has no headless mode, so the setup would be too expensive to run at this scale.
so we built a simulated studio environment that exposes the same exact tool schema that lemonade uses, and implements the roblox studio workspace, so an agent is able to operate just like in production: navigate files, create instances, search roblox docs, and edit scripts. it also includes a renderer built on top of lune, a standalone luau runtime that ships roblox’s real reflection database. the renderer is able to take the local scripts the agent wrote, materializes it into a .rbxm with lune, and rasterizes it to a 1920x1080 png.
this renderer is used both for judging the final completion, but also power a visualCheck tool that the model can call. it renders the current ui and hands the png to a vision model for feedback, which answers in text.
training recipe
dataset. we used ui prompts from lemonade users who opted in to data sharing, each mapped to a correct reference image and a layout family that lemonade has defined in their style guide and has rubrics for.
rewards. we started by defining our rewards based on what lemonade used for their evals: three rewards, one for each eval judge rubric, each normalised to [0, 1] from the judge’s score on the render, plus the fraction of mechanical checks passed. we use gpt 5.6 luna as the judge for cost and latency: each judge sees the render, the prompt, and a reference image lemonade endorsed for that prompt.
algorithm. 8 rollouts per prompt with group-relative advantages, with a dppo update (per-token binary kl masking instead of ratio clipping), kl loss of 0.02 to preserve base capability, and rdpo to balance the reward terms. for hy3 specifically, we trained lora adapters (rank 32) and ran rollouts with high reasoning effort.
validating the recipe on a smaller model
before spinning up a long training job for hy3, we trained qwen 3.5 35b on the recipe first to iterate quickly and validate. this helped ensure simulator and reward correctness. for our final qwen run, we observed mean reward climb from 1.0 to 3.1 (max 4) over the course of 160 updates.
the base 35b model often produced very minimal ui, and didn’t follow instructions to use the provided theme. after training, it was able to improve, but was inconsistent with proper layout and iconography.
base model

step 160

from this, we validated that the recipe can train a model to produce better ui, even if training plateaued, so we moved to a bigger model.
scaling up and tuning the recipe
in training hy3 for this task, we ran into various roadbumps and tuned the recipe accordingly:
ranking rollouts against each other. asking a vision model to score a screenshot from 0 to 100 turned out to be pretty noisy, and the scores bunch up: two renders that a human would rank far apart would often land within a few points of each other, so the model wasn’t getting much signal between an ok ui and a good one. we borrowed from UI2Code^N, which optimizes relative rankings among rendered candidates instead of absolute scores, and added a tournament on top of the absolute rewards. within each rollout group, every pair of renders is sent to the judge head to head along with the reference image, we tally up the wins, and the rank becomes a bonus.
to check that this actually lined up with human preferences, we manually annotated 120 generated candidates (15 groups of 8) to rank from best to worst. absolute scoring agreed with that ranking on 81% of pairs, and the round robin got to 88%, for about 12% more judge calls. it also meant we could use a cheaper judge model: under the tournament it ranked about as well as the expensive ones, at around a tenth of the cost.
adjusting the rewards for correct icon usage. icon usage was only mentioned once across the length judge rubrics, so the reward rarely penalized for picking the wrong ones. as a result, the model learned to reward hack by using the default icons over and over and not searching for more fitting icons.
to address this, we added a separate icon-correctness reward: a fourth judge call that specifically compares icons used in the reference vs the supplied image and scores how many match in quality:
before the icon reward


after the icon reward


in addition to correct icon usage, we also tweaked the existing rubric rewards and mechanical checks to penalize observed gaps such as stretched stud patterns and empty space.
penalizing overlong rollouts. with high reasoning effort enabled, hy3 thinks at length, sometimes past 10k tokens in a single turn before it writes anything. against a 64k training context, that truncated over 40% of rollouts early in training and threw away most of the signal. we adopted the soft overlong punishment from DAPO: full reward up to 49k tokens, scaling linearly down to 0.8x at the 64k hard limit, with truncated rollouts kept in the batch rather than masked out. on the final run, mean response length fell from about 33k to 25k tokens over training, and truncation dropped from 40% of rollouts to 1%.
observations
rewarding form over function: since the reward is purely visual, the model eventually learned to stop trying to implement the actual functionality behind the ui. we audited 240 rollouts across training prompts that asked for real mechanics (a server-backed shop, admin commands, a live leaderboard, and notifications). early in the run, 50% wired up the mechanic. by the end only **6% did, even though reward on those prompts kept going up. the model just learned that a static mockup with hardcoded data scores the same as a working one and is less likely to have issues.
while initially concerning, we observed that the model hadn’t lost the ability to code game mechanics, since the same checkpoint scored 69% on a mechanics-only benchmark vs 70% for the base model. moreover, the harness’s ui system prompt does instructs the model to just implement the ui visually first, before worrying about mechanics. we decided this was in accordance with the product for now: lemonade builds the ui first and ensure it meets what the user expects visually, then wires up game logic in a follow-up turn.
final results
the final hy3 training run went from 1.1 to 4.5 training reward over 113 updates (the max is 5 with the extra reward terms), and held-out reward went from 3.8 at the first eval to 4.4.
at step 0, the base model often didn’t follow the theme/layout fully or just implemented empty panels. by step 108, most rollouts were complete layouts with the correct theming, icons, and layout. for illustration, below are the four highest-reward samples of each rollout group for one training prompt to make a leaderboard, at step 0 and at step 108.
step 0, base hy3




step 108




on lemonade’s eval (39 prompts, 3 attempts each, run through their harness). the rubric score is the mean of the three judges, out of 100:
judge rubric scores, out of 100 | pass rate | ||||||
|---|---|---|---|---|---|---|---|
model | style | layout | quality | average | vs grok | per attempt | vs grok |
grok 4.5 | 79.5 | 79.7 | 76.0 | 78.4 | 44% | ||
hy3 base | 60.1 | 65.3 | 61.3 | 62.2 | -20.7% | 17% | -61.4% |
hy3 after rl | 75.0 | 77.2 | 74.3 | 75.5 | -3.7% | 35% | -20.5% |
the trained model scores well above the base, and lands within a few points of grok 4.5 on average scores alone but doesn’t pass the score threshold as consistently. on cost alone, hy3 comes out to roughly a tenth of the frontier model per token.
future steps
after training the final checkpoint, we helped lemonade deploy the trained model to production, where they offered it to a select cohort of their users as a cheaper option to get feedback. while the model hasn’t quite reached parity with the best frontier model, it proved to be valuable as a much lower cost option. from here, the team at lemonade will be able to take the production traces from opted-in users in their test and further iterate on the training recipe to push training further (for example, filtering dataset examples and forming rewards based on the user telemetry associated with the traces).
want a model trained on your own task? spin one up at app.castform.com