Sapiens2-1B-4K

Sapiens2 is a family of high-resolution vision transformers pretrained on 1 billion human images β€” designed for human-centric tasks such as pose estimation, body-part segmentation, surface normals, and pointmaps.

This repository contains the 1B parameter pretrained backbone, trained at 4K resolution (4096 Γ— 3072, H Γ— W) with a window-tokenizer front-end for tractable token counts. It produces dense per-patch features at 4K input.

Model Details

  • Developed by: Meta
  • Model type: Vision Transformer
  • License: Sapiens2 License
  • Task: pretrain (4K resolution)
  • Format: safetensors
  • File: sapiens2_1b_4k_pretrain.safetensors

Quick Start

Install the Sapiens2 repo (pip install -e .).

Important: This is the 4K variant. You must instantiate with use_tokenizer=True and pass an input tensor of shape (B, 3, 4096, 3072) (H Γ— W).

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from sapiens.backbones.standalone.sapiens2 import Sapiens2

# Build the model and load the 4K pretrained checkpoint
model = Sapiens2(
    arch="sapiens2_1b",
    img_size=(4096, 3072),   # H Γ— W
    patch_size=16,
    use_tokenizer=True,      # required for the 4K variant
).eval().cuda()

ckpt_path = hf_hub_download(
    repo_id="facebook/sapiens2-pretrain-1b-4k",
    filename="sapiens2_1b_4k_pretrain.safetensors",
)
model.load_state_dict(load_file(ckpt_path))

# Forward pass on a single 4K image (RGB; ImageNet normalization recommended)
x = torch.randn(1, 3, 4096, 3072).cuda()
with torch.no_grad():
    features = model(x)[0]  # dense backbone features

Model Card

Field Value
Architecture Sapiens2 ViT (RoPE, GQA, SwiGLU, RMSNorm, QK-norm) + window tokenizer
Backbone parameters 1.607 B
Embedding dim 1536
Layers 40
Attention heads 24
Pretraining resolution 4096 Γ— 3072 (H Γ— W)
Patch size 16
Window size 4
Pretraining data 1B human images

Sapiens2 Family

Model Params FLOPs Embed dim Layers Heads
Sapiens2-0.1B 0.114 B 0.342 T 768 12 12
Sapiens2-0.4B 0.398 B 1.260 T 1024 24 16
Sapiens2-0.8B 0.818 B 2.592 T 1280 32 16
Sapiens2-1B 1.462 B 4.715 T 1536 40 24
Sapiens2-1B-4K (this) 1.607 B β€” 1536 40 24
Sapiens2-5B 5.071 B 15.722 T 2432 56 32

See the Sapiens2 Collection for all variants and downstream task checkpoints (pose, segmentation, normals, pointmaps).

Intended Use

  • Feature extraction for human-centric downstream tasks at 4K resolution
  • Initialization for fine-tuning high-resolution task heads (pose, segmentation, normals, pointmap)
  • Research on human-centric vision at high resolution

License

Released under the Sapiens2 License.

Citation

@article{khirodkarsapiens2,
  title={Sapiens2},
  author={Khirodkar, Rawal and Wen, He and Martinez, Julieta and Dong, Yuan and Su, Zhaoen and Saito, Shunsuke},
  journal={arXiv preprint arXiv:2604.21681},
  year={2026}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including facebook/sapiens2-pretrain-1b-4k

Paper for facebook/sapiens2-pretrain-1b-4k