Context
Local image generation presents a different problem from text inference: large models, significant memory requirements and a direct relationship between resolution, generation time and available resources.
This experiment studies how far Apple Silicon can go when running generative image models entirely locally.
Goal
Compare different models and inference pipelines to understand their practical limits and determine which approaches are best suited to each workload.
The goal is not simply to generate an image successfully. We also want to understand:
- generation time
- resource usage
- viable resolutions
- stability
- prompt fidelity
- visual consistency
- behaviour with complex scenes
Experimentation
We start from a baseline model that provides a known reference and compare it against new alternatives running locally.
The goal is not to replace a model simply because another one is newer, but to determine whether it provides meaningful improvements under our execution conditions.
Prompt
↓
Prompt processing
↓
Image model
↓
Local inference
↓
Image
↓
Evaluation
Each generation therefore becomes an observation about system behaviour rather than just a visual result.
Limits
The experiments show that increasing resolution does not scale indefinitely.
As pixel count increases, so do memory pressure, inference time and the likelihood of reaching practical limits in either the hardware or the pipeline.
This makes maximum stable resolution an experimental property of the system rather than simply a configurable parameter.
Comparison
The baseline model provides a known reference against which new alternatives can be evaluated.
Tests compare the different approaches under similar conditions, observing not only speed and quality but also structural errors, prompt adherence, variation between generations and behaviour with difficult compositions.
The results of these experiments determine whether an alternative provides enough benefit to justify changes to the pipeline.
Status
Active.
Testing continues while we compare different models and local inference configurations on Apple Silicon.