Image2LEGO starts with the most intuitive input: a photo or reference image. The hard part is translating pixels into a model whose silhouette still looks right while every brick remains real, connected, supported, and buildable.
The first stage uses multimodal vision to identify the subject, foreground silhouette, depth cues, semantic parts, symmetry, major colors, and intended scale. The system then creates a simplified 3D plan before mapping volumes and surfaces to real brick types.
BrickGPT and BrickNet suggest two research directions for constrained generation: physical feedback during brick placement, and graph-based representation of part connections. A production image-to-brick pipeline would still need its own geometry and connection validation before instructions or exports could be trusted.
To try the current product flow, see the photo-to-bricks guide. The research systems discussed here are not presented as implemented features of Image2LEGO.
Generate, test, roll back
Predict the next brick, reject invalid choices, and retreat when the structure becomes unsafe.
Generate through connections
Position parts through connection relationships rather than relying only on raw 3D coordinates.
BrickGPT: physical feedback inside generation
The first May 2025 preprint used the name LegoGPT and the title Generating Physically Stable and Buildable LEGO Designs from Text. The latest 2025 revision and official project page use BrickGPT and the title Generating Physically Stable and Buildable Brick Structures from Text.
The system fine-tunes an autoregressive language model to predict a sequence of brick placements. The critical improvement is not language alone. During inference, validity checks prune infeasible placements, while physics-aware rollback lets the system retreat when a sequence produces an unstable design.
next brick
connection
roll back
The current paper calls the released dataset StableText2Brick and reports more than 47,000 structures derived from over 28,000 unique 3D objects. The reusable product lesson is straightforward: let a generative model propose, but let constraint systems decide.
“we employ an efficient validity check and physics-aware rollback”
— Ava Pun et al., BrickGPT (2025). Quoted from the paper's abstract.
BrickNet: make connectivity the language
BrickNet starts from a different observation: raw coordinates are a fragile way to describe assemblies. What defines a brick model is the network of studs, tubes, pins, axles, and other connection semantics between parts.
Its graph-backed program representation models those relationships directly. The 2026 paper reports 320,808 samples, 9,743 part variants, and 40,549,969 placed brick instances in its LDraw dataset—a much broader vocabulary than voxel-like systems limited to a few basic bricks.
“it is the spatial relationships of the parts which define the whole”
— Peter Kulits and Cordelia Schmid, BrickNet (2026). Quoted from the paper's abstract.
Store topology separately from appearance. A connection graph can support placement checks, step ordering, sub-build discovery, and part substitution.
The image-to-brick conversion pipeline
The core task is converting visual evidence into a buildable representation. Neither a general-purpose model nor a specialist research model should own the entire workflow; each layer needs one clear job.
Read the image
Separate the subject from its background and extract silhouette, depth cues, semantic parts, symmetry, colors, and scale.
Build a 3D shape plan
Convert 2D evidence into approximate volumes, surfaces, proportions, sub-builds, and structural anchors.
Map shapes to real bricks
Select available part types and propose connected placements while preserving recognizable visual features.
Validate and repair
Detect collisions, floating parts, weak supports, unreachable placements, and unstable step order.
Preview, edit, and export
Render the 3D result, let users inspect it, create building steps and parts lists, and export LDraw, CSV, or JSON.
Where current OpenAI and Claude models fit
As of October 2026, frontier general-purpose models are most useful as planners, critics, and tool users—not as substitutes for geometry and physics.
A sensible orchestration starts with a cost-efficient model and escalates only when uncertainty or validation failures justify it. The specialist geometry layer remains deterministic and testable regardless of which general model is selected.
Research prototype is not a production guarantee
A visually plausible model may still use rare parts, create fragile assemblies, or produce steps that are difficult for human hands. Every export should expose validation status and let users inspect the result before buying parts.
Image2LEGO is inspired by these research directions, but this article does not claim that the website implements every capability or reproduces every published result. The practical objective is a transparent workflow where users can inspect models, parts, build steps, and exports.
Primary sources
Frequently asked questions
How does an image become a brick model?
A vision model first extracts recognizable shapes, colors, depth cues, and semantic parts. A planning layer approximates those features in 3D, then a brick engine maps them to real parts and validates connections before producing steps and exports.
Can a general AI model generate a finished brick file by itself?
It can understand the image, plan the model, and emit structured candidates, but reliable output still needs part catalogs, connection rules, collision checks, and structural validation.
What is LDraw?
LDraw is an open file format and parts ecosystem used to describe virtual brick models, scenes, and building steps.
Which model should a production pipeline start with?
Start with a balanced vision-capable model, measure it on your own reference images, and reserve frontier tiers for difficult planning or repair cases.