Skip to content

Repository files navigation

RPA: Scalable Black-Box Model Attribution for Images

Asaf Livne  ·  Amir Jevnisek  ·  Shai Avidan

Tel Aviv University

arXiv License Python PyTorch

Attribute an image to the model that generated it — from raw pixels, with a ~6M-parameter CNN.

One prompt rendered by twenty-five text-to-image models

Same prompt, twenty-five generators. Telling them apart is the attribution problem.


📰 News

  • [2026-09] Checkpoints, configs and inference updated to the 64×64-patch model of the current paper version.
  • [2026-08] Code and pretrained checkpoints released.
  • [2026-08] arXiv preprint released.

🧠 Overview

The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? Existing methods have grown as elaborate as the generators they target, on the assumption that a more sophisticated generator demands a more sophisticated attributor. We show it does not.

RPA (Raw-Patch Attribution) attributes images in the strictest black-box setting with a lightweight CNN. Despite its simplicity it attributes more models at higher accuracy than prior work, is data-efficient, runs at a cost independent of the number of candidate models, and stays robust to the compression, blur, and resizing images undergo in the wild.

Training for closed-set attribution also yields a versatile feature extractor: the same representation recovers model lineage without supervision, flags unseen generators, and admits new models through few-shot adaptation rather than retraining.

Why RPA?

  • 🪶 Small. ~6M parameters, no pretrained encoder, no foundation-model backbone.
  • 🔒 Strictly black-box. Image only — no generator weights, no VAE, no prompts.
  • 📐 Resolution-invariant. 64² to 4096² without retraining; cost is independent of the number of candidate models.
  • 🧩 More than a classifier. The same features give open-set rejection, unsupervised lineage, and few-shot adaptation for free.

🔍 Method

Three steps, one small network.

Step What happens
01 Patch division Split the image into fixed 64×64 RGB patches, overlapping at the edges so every pixel is covered exactly once in the weighting.
02 Per-patch classification A compact ~6M-parameter CNN — three strided-conv blocks, global average pool, linear head — scores each patch against every candidate generator in one forward pass.
03 Aggregation Per-patch scores combine into one image-level label via an overlap-corrected weighted average (logit_avg).

📊 Results

Black-box attribution

From the image alone, with no model access at all. Top-1 accuracy (%) and per-image inference time (ms, mean ± std, RTX 5090).

Method Params (M) Infer. (ms) DRAGON (25) OpenFake (27)
DE-FAKE 151 9.96 ± 1.31 62.0 57.3
USIA 427 9.61 ± 1.01 52.1 50.1
LIDA 23.5 23.6 ± 12.5 24.0 17.7
OCC-CLIP 151 7.87 ± 1.12 8.6 15.8
RPA (ours) 5.9 2.79 ± 0.04 98.9 95.0

EfficientFormer releases neither code nor data; it reports 91.0 % on 13 classes of private data with ~4× more images per class.

White-box comparison

On AEDR's eight-model benchmark, where competitors are given each candidate's autoencoder. Mean pairwise accuracy and per-image inference time.

Method Access Infer. (s) Acc. (%)
LatentTracer Model weights 24.06 ± 0.01 70.3
AEDR VAE weights 0.267 ± 0.003 95.1
RPA (ours) Image only 0.00279 ± 0.00004 99.5

Two orders of magnitude faster than the closest competitor, and strictly black-box; 98.8 % in the harder 8-way single-label setting. Baseline accuracies are taken from AEDR; all inference times are re-measured on our hardware, per candidate model (ours: one fp16 pass over the 256 patches of a 1024² query, PNG decode excluded).

Open-set & discovery

Trained on 17 of OpenFake's 27 generators, with 10 held out as unseen:

Task Metric
Rejecting unseen generators AU-OSCR 0.875 ± 0.037 over five draws (best draw 0.918)
Clustering the 10 unseen sources ARI 0.64 · NMI 0.84 · 98 % purity (7 clusters recovered against 10 true sources; five-draw mean ARI 0.54 · NMI 0.78 · 93 % purity)
Recovering model lineage (OpenFake-27, no supervision) cophenetic r = 0.910

Adaptation

A frozen 17-class OpenFake backbone with a freshly fit 27-way linear head (28 K parameters, ~2 min). Mean ± std over five 17/10 draws.

Eval. subset Before After
Original 17 95.8 ± 1.3 93.3 ± 1.8
New 10 — 85.1 ± 4.2
All 27 — 90.3 ± 0.5

Across benchmarks, an OpenFake backbone with a new head attributes DRAGON's 25 generators at 96.6 % (vs. 98.9 % for DRAGON's own model); the reverse reaches 86.8 % (vs. 95.0 %).

Few-shot adaptation

Nine unseen GenImage generators, N labeled images per class, LIDA's protocol. Only a linear head is fit — the backbone stays frozen.

Method 1-shot 10-shot
ResNet 17.4 21.4
DIRE 14.3 17.2
ESSP 17.0 22.4
LIDA 40.4 54.0
RPA (ours, OpenFake backbone) 47.7 ± 3.5 72.0 ± 0.8
RPA (ours, DRAGON backbone) 52.0 ± 3.6 72.4 ± 0.9

Baselines as published in LIDA.


🛠️ Installation

git clone https://github.com/Asaf-Livne/raw-patch-attribution.git
cd raw-patch-attribution
pip install -e .

Optional extras, each pulling in only what its entry point needs:

pip install -e ".[cluster]"    # hdbscan + umap-learn, for `cluster.py discovery`
pip install -e ".[download]"   # datasets, for `scripts/download_data.py`
pip install -e ".[plot]"       # matplotlib, for training-curve PNGs

🎮 Pretrained checkpoints

Both headline classifiers are published as release assets (~25 MB each), keeping the repository light.

Checkpoint Benchmark Classes Top-1 Download
dragon_25class.pt DRAGON 25 98.9 % ⬇
openfake_27class.pt OpenFake 27 95.0 % ⬇
BASE=https://github.com/Asaf-Livne/raw-patch-attribution/releases/download/v1.0
curl -L -o checkpoints/dragon_25class.pt   $BASE/dragon_25class.pt
curl -L -o checkpoints/openfake_27class.pt $BASE/openfake_27class.pt
shasum -a 256 -c checkpoints/SHA256SUMS

See checkpoints/README.md for loading them directly in Python.


📦 Data preparation

python scripts/download_data.py dragon   --output-dir datasets --config Regular
python scripts/download_data.py openfake --output-dir datasets/openfake --samples-per-class 1400

DRAGON downloads pre-split. OpenFake downloads flat per class, and PatchFolderDataset applies a seeded 750/250/400 train/val/test split at load time. GenImage — used only as the few-shot adaptation target, as a 9-generator subset under LIDA's protocol — is not included here; prepare it per that benchmark's own release.

Point each config's dataset.data_root at your data. Config strings expand ${ENV_VAR} and ~, so the shipped configs stay machine-independent:

export DRAGON_DATA_ROOT=datasets/dragon_regular
export OPENFAKE_DATA_ROOT=datasets/openfake

🚀 Training

python train.py --config configs/dragon_25class.yaml --output runs/dragon_25class --seed 42

Swap in configs/openfake_27class.yaml for OpenFake. A run directory holds best.pt, last.pt, a config snapshot, curves.csv, and channel statistics; re-running against the same --output auto-resumes from last.pt.

For the robustness variant, train with on-the-fly JPEG / blur / resize corruption — same code path, different config:

python train.py --config configs/dragon_20class_robust.yaml --output runs/dragon_robust --seed 42

🎯 Evaluation

python eval.py --checkpoint checkpoints/dragon_25class.pt \
    --config configs/dragon_25class.yaml \
    --split test --aggregation logit_avg --num_patches "1 16 256"

--num_patches accepts a list; each image is scored with min(N, its own tile count) patches, so a budget at or above the tile count means all patches (256 for a 1024² image at 64×64) and mixed-resolution benchmarks are handled correctly. Results land in summary.json plus per-budget confusion matrices.

Inference tiles the image on the GPU in one gather and runs the BatchNorm-folded network in fp16 (fp32 on CPU): about 3 ms per 1024² image (256 patches) on an RTX 5090, PNG decode excluded, with no change in accuracy.

For the "what does the CNN see" frequency analysis, set dataset.input_filter to lowpass, highpass, or fftmag in the eval config.


🔬 Beyond closed-set attribution

The same trained backbone drives three further capabilities — no retraining.

Open-set attribution — reject generators never seen in training

Carve out a known-class subset, train on it, then score against the full label map: every class the checkpoint was not trained on is treated as unseen. Detection uses max-softmax-probability under all-patch logit averaging, and is reported threshold-free (AUROC, AU-OSCR).

python scripts/split_known_unknown.py --label-map configs/label_maps/openfake_27class.json \
    --n-known 17 --seed 42 --output configs/label_maps/openfake_17known_seed42.json

python train.py --config configs/openfake_17known.yaml --output runs/openset_seed42 --seed 42

python openset.py --checkpoint runs/openset_seed42/best.pt \
    --config configs/openfake_27class.yaml --output runs/openset_seed42/eval
Lineage & unseen-source discovery — unsupervised structure in the features

lineage clusters per-generator mean features hierarchically and reports the cophenetic correlation; discovery runs UMAP + HDBSCAN over the classes the checkpoint was never trained on and scores the result against their true labels (ARI / NMI / purity). Neither uses any lineage or discovery labels.

python cluster.py lineage --checkpoint checkpoints/openfake_27class.pt \
    --config configs/openfake_27class.yaml --output runs/openfake_27class/lineage

python cluster.py discovery --checkpoint runs/openset_seed42/best.pt \
    --config configs/openfake_27class.yaml --output runs/openset_seed42/discovery

Requires the cluster extra.

Adaptation — admit new generators by fitting a linear head

The backbone is frozen and its penultimate features cached once, so fitting the head takes seconds to minutes instead of a full retrain. All three published regimes share one script; they differ only in what target data feeds the cache.

# full: extend a known-subset backbone to the benchmark's complete class set
python adapt.py --checkpoint runs/openset_seed42/best.pt \
    --config configs/openfake_27class.yaml --output runs/adapt_full27 --transplant

# cross-dataset: freeze a backbone trained on one benchmark, fit a head on another
python adapt.py --checkpoint checkpoints/dragon_25class.pt \
    --config configs/openfake_27class.yaml --output runs/cross_d2o

# few-shot on GenImage-9 (LIDA's protocol)
python adapt.py --checkpoint checkpoints/openfake_27class.pt \
    --config configs/genimage_9class.yaml --output runs/fewshot_10shot \
    --shots 10 --train_views 20 --val_split val --test_split val

--transplant keeps the trained head row for any class shared by name with the source checkpoint, so only genuinely new classes start from scratch.


📁 Repository layout

train.py       train a classifier (on-the-fly augmentation)
eval.py        closed-set evaluation: multi-patch aggregation, robustness, frequency analysis
openset.py     open-set attribution (reject unseen generators)
cluster.py     lineage recovery + unseen-source clustering
adapt.py       adapt a frozen backbone to new classes (full / cross-dataset / few-shot)

datalib/       dataset, patch tiling, augmentation, open-set metrics
models/        the CNN
configs/       one YAML per run, plus label_maps/ ({class_name: folder}, one per benchmark)
scripts/       data download + known/unknown split generator
checkpoints/   download target for the pretrained classifiers

Every entry point takes --config <path> plus its own flags.


📝 Citation

@misc{livne2026scalableblackboxmodelattribution,
  title         = {Scalable Black-Box Model Attribution for Images},
  author        = {Asaf Livne and Amir Jevnisek and Shai Avidan},
  year          = {2026},
  eprint        = {2608.15652},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2608.15652}
}

💡 Acknowledgments

This work builds on the following public benchmarks and baselines:

  • DRAGON — 25-generator attribution benchmark.
  • OpenFake — 27-generator benchmark.
  • GenImage / LIDA — few-shot adaptation protocol.
  • AEDR — the white-box eight-model comparison benchmark.

📄 License

Released under the MIT License.

About

Attribute an image to the model that generated it — from raw pixels, with a ~6M-parameter CNN (RPA).

Topics

Resources

Stars

35 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages