Qwen-Image 2.1 Edit: equirectangular 360 panorama LoRA

An edit LoRA for Qwen-Image 2.1. Give it 1-3 ordinary photos of a place, taken from one spot while turning the camera, plus a one-line description of the scene. It returns a full 2:1 equirectangular 360 panorama of that place, ready for a 360 viewer, a VR photo or a skybox / environment map.

The LoRA places the views where they belong on the sphere and invents the rest of the surroundings (behind the camera, the sky overhead, the ground below) to match. The first image always lands at the center of the panorama, looking straight ahead.

Example

One input view, with the scene description "A grassy hilltop at dawn with dry tussock grass and scattered rocks, overlooking a lake and dam among flat-topped distant mountains, soft warm low-contrast sunrise light under a clear sky."

Input (<image1>)
input view

Output, 1536x768, seed 7:

generated panorama

The center of the panorama, projected back to a perspective view with the input's camera angle (pitched 17° down, 86° field of view). The lake, the hills and both groups of rocks come back where they were in the input, softer because 1536 px have to cover the full 360°:

center of the panorama re-projected to the input camera

The input is a crop of the CC0 Poly Haven qwantani_dawn panorama, which is in the training set. In training this crop was the second of three views, about 171° away from the center; here it is the only view and becomes the center.

How it was trained

Each training sample is built from a real 360 panorama (CC0 HDRIs from Poly Haven, tonemapped):

  1. Pick 1-3 random perspective views inside the panorama: random direction, a slight tilt up or down, a field of view between 40° and 105°, and an aspect ratio of 1:1, 4:3, 3:2, 16:9, 3:4 or 2:3. Views that are almost blank (plain wall, empty sky) are rejected, and views in the same sample don't overlap much.
  2. Render each view as a normal flat photo from the 8K source.
  3. Rotate the panorama so the first view sits at its center; that rotated panorama is the target.
  4. Caption it Transform this set of images into an equirectangular 360 panorama. Scene: <description>.

The previews below show two training samples. The top half is the target panorama with each view's footprint outlined; the bottom half shows the views exactly as the model receives them, with the caption underneath.

One view. <image1> is a square crop of the steps of Rhodes Memorial. The target panorama has it at the center and the rest of the terrace, the city and the sky around it:

training sample with one view

Three views. <image1> (portrait) looks down the road from an overpass and sits at the center. <image2> looks about 80° to the right, <image3> about 100° to the left. The parts between and behind them are what the model learns to fill in:

training sample with three views

Outlines near the top and bottom look curved because straight lines bend in equirectangular projection, the same way the railing bends in the target.

Training data: 232 panoramas (2 samples each, 464 total: 161 with one view, 156 with two, 147 with three), plus 8 held out for validation. Categories: streets and towns, interiors, fields and countryside, mountains and hills, parks and gardens, coast and water, abandoned places and ruins, forests, deserts, industrial sites and a few studios.

What input to give it

  • 1-3 photos taken from the same spot, turning the camera between shots (like the first steps of shooting a panorama). Photos of the same place from different positions, or of different places, don't fit the model's assumption of a single viewpoint.
  • The most important view first. <image1> becomes the center of the panorama. The other views are placed around it by their content, so their order and exact angle don't need to be given.
  • Ordinary photos, roughly level. Normal lenses, about 40-105° wide, any common aspect ratio (landscape, portrait or square). Training views were tilted at most 8° (first view) or 20° (others), so straight-up sky or straight-down floor shots are outside what it learned. Fisheye or already-stitched panoramas are not expected input.
  • Views with something in them. Texture, structure and landmarks help; a blank wall or plain sky gives the model little to place.
  • More views, more faithful. Whatever no view shows is invented. Views spread around the full circle give a panorama closer to the real place.
  • A short scene description as the caption's Scene: part (below).

Prompt

Transform this set of images into an equirectangular 360 panorama. Scene: <description>

The description is one sentence (about 15-35 words) in the style of the training captions: the place and its main features, then the lighting (time of day, weather, natural or artificial, contrast), then optionally the mood. Describe the whole surrounding space, not just what the views show. No <image1> references or edit instructions are needed. Training caption examples:

  • A small empty house interior with warm morning light through door and windows, tiled floor and soft mid-contrast shadows.
  • Dim industrial warehouse lighting with strong warm artificial overhead lamps and skylights, high-contrast, gritty utilitarian mood.
  • A frozen winter lake at sunrise, soft low-contrast natural light with cool blue snow and a warm glow in partly cloudy skies.

Use (ComfyUI)

  • LoraLoaderModelOnly on the Qwen-Image 2.1 diffusion model, strength 1.0.
  • TextEncodeQwenImage21 with the views as image_1..image_3 and resolution 1088 (references then match their training size).
  • Sample from an empty 1536x768 latent, not the encoder's latent output (that one follows <image1>'s aspect ratio), 25 steps, CFG 1.0, euler / simple. Larger 2:1 sizes such as 2048x1024 also work and come out sharper.

Limitations

  • Regions no view covers are plausible inventions, not reconstructions.
  • The left and right edges are not forced to match, so the seam behind the viewer may not wrap perfectly.
  • Trained without other LoRAs; stacking a style or realism LoRA may weaken the projection.

Training details

  • Trainer: ai-toolkit (qwen_image_2), rank 32, 3000 steps, lr 1e-4, batch 1, AdamW 8-bit, bf16 transformer, caption dropout 0.05. About 4.25 h on one RTX PRO 6000.
  • Targets at 1536x768; references scaled to the same pixel area.
  • Weights are in ComfyUI key format.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Gogodr/qwen-image-2.1-edit-pano360-lora

Adapter
(91)
this model