Velgrina AI Remodel: Separating Photo Understanding from Image Generation

Inside Velgrina’s AI Remodel feature: kitchen-photo understanding, image generation, daily limits and an administrative provider switch.

Velgrina’s AI Remodel feature starts with a kitchen photo and produces a new visualization. I implemented it as two different AI tasks: GPT-4o reads the photograph, then Gemini or gpt-image-1 produces the visual result. The feature includes a daily limit and an administrative switch for the generation provider.

This was a useful contrast with the store’s product-help chat. A visual feature has a different output, a different failure experience and a greater risk that a convincing result will be mistaken for something physically exact.

Understanding and rendering have different jobs

The analysis stage interprets the uploaded image. The rendering stage creates a new image from that context and the requested transformation. Separating them makes the responsibilities easier to reason about, even though both stages involve models.

For a kitchen, the distinction is practical. Recognizing a cabinet, a window or a worktop provides context for a visualization. It does not establish the room’s exact dimensions, reveal hidden plumbing or prove that a particular product will fit. A rendered image is therefore an exploration of an idea, rather than a measured renovation specification.

That limitation should shape the experience. I want users to understand what they are receiving before they interpret the result. A visually attractive output should not imply an installation guarantee or replace an actual product measurement.

The handover is an interface

When one model’s output becomes another model’s input, the handover deserves the same care as an ordinary service boundary. The relevant context should be clear enough to inspect: what was observed, what transformation was requested and which details are intended to remain recognizable.

For this class of feature, I review examples where the original image is dark, cluttered or taken from an unusual angle. I also compare the result with the input for structural drift. If a window disappears or the room’s geometry changes, the visualization may still be attractive while becoming less useful to the customer.

These are evaluation criteria, not a claim that the system enforces exact geometric preservation. The distinction is important because image quality and task fidelity can move in different directions.

A provider switch creates operational choices

The administrative provider switch was part of the implementation. It gave the application an explicit place to select the rendering backend instead of coupling that decision to a single permanent integration.

An alternative provider is still a different execution path. Output style, response format, error behavior and processing time can vary. I treat a switch as a change that needs review with representative inputs, rather than assuming two image APIs are interchangeable because they both return an image.

The same reasoning applies to diagnostics. A useful failure report should make clear whether the upload, interpretation or rendering stage failed. That allows the operator to investigate the right boundary and helps avoid an unnecessary repeat of work that already succeeded.

Daily limits belong in the experience

AI Remodel includes a daily usage limit. For an image-generation feature, the limit is part of both resource control and customer communication. A user should know whether another request is available and what a rejected attempt means.

My review questions include the quota boundary, repeated submissions and interrupted requests. The implementation policy needs to define when an attempt consumes allowance and how a failed attempt is handled. The interface should then describe that policy consistently. An unexplained “try again” message is especially unhelpful when the next attempt may consume the last available request.

This project expanded my AI work from text responses into a customer-facing multimodal workflow. The core engineering challenge was coordinating interpretation, generation, operational controls and user expectations. The resulting image is the visible part; the product needs those surrounding decisions to make it understandable and maintainable.

Updated 26 September 2026.