← Lab
VJML / LAB
2026.10 ACTIVE

Local image generation pipeline

Local AI / Infrastructure

Context

Generating an image is only one part of the system.

ai.vjml.es uses local infrastructure to transform a request into an image through an automated pipeline. Several stages around the generation model have a significant influence on the final result.

This work explores how to design, simplify and improve that pipeline without relying on external services for inference.

Architecture

At a high level, the system can be represented as:

User input
    ↓
Prompt processing
    ↓
Generation request
    ↓
Local inference
    ↓
Image
    ↓
Delivery

Each of these stages can evolve independently.

Separating the interface, prompt processing and generation engine makes it possible to experiment with different approaches without coupling the whole system to a specific implementation.

Prompt pipeline

An important part of the work happens before generation.

User input needs to be transformed into sufficiently structured instructions for the system to interpret intent, composition, style and other relevant aspects of the image.

The goal is not simply to produce longer prompts, but to make the transformation consistent and ensure that each part of the pipeline has a clear purpose.

Evolution

The pipeline has gone through successive iterations as new test cases reveal different behaviours.

Each iteration helps expose problems that are not always obvious from isolated generations:

  • duplicated elements
  • incorrect anatomy
  • loss of intent
  • inconsistent composition
  • insufficient variation
  • instructions that are present but have little practical effect

This turns pipeline development into an iterative process of generation, observation and correction.

Generation engine separation

The generation model is treated as an interchangeable component.

The system should be able to compare the behaviour of the baseline model against new alternatives without making the rest of the architecture depend on the specific identity of the engine being used.

             ┌─ Generation engine A
Pipeline ────┤
             └─ Generation engine B
                     ↓
                  Compare

This makes it possible to evaluate new options while keeping the rest of the system stable.

Local infrastructure

Generation runs on infrastructure under our control.

In addition to keeping inference local, this architecture makes it possible to directly observe performance, memory limits, viable resolutions and system behaviour under different workloads.

The frontend and generation infrastructure can run on separate systems while operating as components of the same service.

Status

Active.

We are currently refining prompt processing and comparing system behaviour across different generation configurations.