Asaf Livne · Amir Jevnisek · Shai Avidan
Tel Aviv University
Attribute an image to the model that generated it — from raw pixels, with a ~6M-parameter CNN.
Same prompt, twenty-five generators. Telling them apart is the attribution problem.
- [2026-09] Checkpoints, configs and inference updated to the 64×64-patch model of the current paper version.
- [2026-08] Code and pretrained checkpoints released.
- [2026-08] arXiv preprint released.
The rapid proliferation of generative models raises the model attribution problem: given only an image, can we determine which model produced it? Existing methods have grown as elaborate as the generators they target, on the assumption that a more sophisticated generator demands a more sophisticated attributor. We show it does not.
RPA (Raw-Patch Attribution) attributes images in the strictest black-box setting with a lightweight CNN. Despite its simplicity it attributes more models at higher accuracy than prior work, is data-efficient, runs at a cost independent of the number of candidate models, and stays robust to the compression, blur, and resizing images undergo in the wild.
Training for closed-set attribution also yields a versatile feature extractor: the same representation recovers model lineage without supervision, flags unseen generators, and admits new models through few-shot adaptation rather than retraining.
- 🪶 Small. ~6M parameters, no pretrained encoder, no foundation-model backbone.
- 🔒 Strictly black-box. Image only — no generator weights, no VAE, no prompts.
- 📐 Resolution-invariant. 64² to 4096² without retraining; cost is independent of the number of candidate models.
- 🧩 More than a classifier. The same features give open-set rejection, unsupervised lineage, and few-shot adaptation for free.
Three steps, one small network.
| Step | What happens | |
|---|---|---|
| 01 | Patch division | Split the image into fixed 64×64 RGB patches, overlapping at the edges so every pixel is covered exactly once in the weighting. |
| 02 | Per-patch classification | A compact ~6M-parameter CNN — three strided-conv blocks, global average pool, linear head — scores each patch against every candidate generator in one forward pass. |
| 03 | Aggregation | Per-patch scores combine into one image-level label via an overlap-corrected weighted average (logit_avg). |
From the image alone, with no model access at all. Top-1 accuracy (%) and per-image inference time (ms, mean ± std, RTX 5090).
| Method | Params (M) | Infer. (ms) | DRAGON (25) | OpenFake (27) |
|---|---|---|---|---|
| DE-FAKE | 151 | 9.96 ± 1.31 | 62.0 | 57.3 |
| USIA | 427 | 9.61 ± 1.01 | 52.1 | 50.1 |
| LIDA | 23.5 | 23.6 ± 12.5 | 24.0 | 17.7 |
| OCC-CLIP | 151 | 7.87 ± 1.12 | 8.6 | 15.8 |
| RPA (ours) | 5.9 | 2.79 ± 0.04 | 98.9 | 95.0 |
EfficientFormer releases neither code nor data; it reports 91.0 % on 13 classes of private data with ~4× more images per class.
On AEDR's eight-model benchmark, where competitors are given each candidate's autoencoder. Mean pairwise accuracy and per-image inference time.
| Method | Access | Infer. (s) | Acc. (%) |
|---|---|---|---|
| LatentTracer | Model weights | 24.06 ± 0.01 | 70.3 |
| AEDR | VAE weights | 0.267 ± 0.003 | 95.1 |
| RPA (ours) | Image only | 0.00279 ± 0.00004 | 99.5 |
Two orders of magnitude faster than the closest competitor, and strictly black-box; 98.8 % in the harder 8-way single-label setting. Baseline accuracies are taken from AEDR; all inference times are re-measured on our hardware, per candidate model (ours: one fp16 pass over the 256 patches of a 1024² query, PNG decode excluded).
Trained on 17 of OpenFake's 27 generators, with 10 held out as unseen:
| Task | Metric |
|---|---|
| Rejecting unseen generators | AU-OSCR 0.875 ± 0.037 over five draws (best draw 0.918) |
| Clustering the 10 unseen sources | ARI 0.64 · NMI 0.84 · 98 % purity (7 clusters recovered against 10 true sources; five-draw mean ARI 0.54 · NMI 0.78 · 93 % purity) |
| Recovering model lineage (OpenFake-27, no supervision) | cophenetic r = 0.910 |
A frozen 17-class OpenFake backbone with a freshly fit 27-way linear head (28 K parameters, ~2 min). Mean ± std over five 17/10 draws.
| Eval. subset | Before | After |
|---|---|---|
| Original 17 | 95.8 ± 1.3 | 93.3 ± 1.8 |
| New 10 | — | 85.1 ± 4.2 |
| All 27 | — | 90.3 ± 0.5 |
Across benchmarks, an OpenFake backbone with a new head attributes DRAGON's 25 generators at 96.6 % (vs. 98.9 % for DRAGON's own model); the reverse reaches 86.8 % (vs. 95.0 %).
Nine unseen GenImage generators, N labeled images per class, LIDA's protocol. Only a linear head is fit — the backbone stays frozen.
| Method | 1-shot | 10-shot |
|---|---|---|
| ResNet | 17.4 | 21.4 |
| DIRE | 14.3 | 17.2 |
| ESSP | 17.0 | 22.4 |
| LIDA | 40.4 | 54.0 |
| RPA (ours, OpenFake backbone) | 47.7 ± 3.5 | 72.0 ± 0.8 |
| RPA (ours, DRAGON backbone) | 52.0 ± 3.6 | 72.4 ± 0.9 |
Baselines as published in LIDA.
git clone https://github.com/Asaf-Livne/raw-patch-attribution.git
cd raw-patch-attribution
pip install -e .Optional extras, each pulling in only what its entry point needs:
pip install -e ".[cluster]" # hdbscan + umap-learn, for `cluster.py discovery`
pip install -e ".[download]" # datasets, for `scripts/download_data.py`
pip install -e ".[plot]" # matplotlib, for training-curve PNGsBoth headline classifiers are published as release assets (~25 MB each), keeping the repository light.
| Checkpoint | Benchmark | Classes | Top-1 | Download |
|---|---|---|---|---|
dragon_25class.pt |
DRAGON | 25 | 98.9 % | ⬇ |
openfake_27class.pt |
OpenFake | 27 | 95.0 % | ⬇ |
BASE=https://github.com/Asaf-Livne/raw-patch-attribution/releases/download/v1.0
curl -L -o checkpoints/dragon_25class.pt $BASE/dragon_25class.pt
curl -L -o checkpoints/openfake_27class.pt $BASE/openfake_27class.pt
shasum -a 256 -c checkpoints/SHA256SUMSSee checkpoints/README.md for loading them directly in Python.
python scripts/download_data.py dragon --output-dir datasets --config Regular
python scripts/download_data.py openfake --output-dir datasets/openfake --samples-per-class 1400DRAGON downloads pre-split. OpenFake downloads flat per class, and
PatchFolderDataset applies a seeded 750/250/400 train/val/test split at load
time. GenImage — used only as the few-shot adaptation target, as a
9-generator subset under LIDA's protocol — is not included here; prepare it
per that benchmark's own release.
Point each config's dataset.data_root at your data. Config strings expand
${ENV_VAR} and ~, so the shipped configs stay machine-independent:
export DRAGON_DATA_ROOT=datasets/dragon_regular
export OPENFAKE_DATA_ROOT=datasets/openfakepython train.py --config configs/dragon_25class.yaml --output runs/dragon_25class --seed 42Swap in configs/openfake_27class.yaml for OpenFake. A run directory holds
best.pt, last.pt, a config snapshot, curves.csv, and channel statistics;
re-running against the same --output auto-resumes from last.pt.
For the robustness variant, train with on-the-fly JPEG / blur / resize corruption — same code path, different config:
python train.py --config configs/dragon_20class_robust.yaml --output runs/dragon_robust --seed 42python eval.py --checkpoint checkpoints/dragon_25class.pt \
--config configs/dragon_25class.yaml \
--split test --aggregation logit_avg --num_patches "1 16 256"--num_patches accepts a list; each image is scored with min(N, its own tile
count) patches, so a budget at or above the tile count means all patches (256
for a 1024² image at 64×64) and mixed-resolution benchmarks are handled
correctly. Results land in summary.json plus per-budget confusion matrices.
Inference tiles the image on the GPU in one gather and runs the BatchNorm-folded network in fp16 (fp32 on CPU): about 3 ms per 1024² image (256 patches) on an RTX 5090, PNG decode excluded, with no change in accuracy.
For the "what does the CNN see" frequency analysis, set
dataset.input_filter to lowpass, highpass, or fftmag in the eval config.
The same trained backbone drives three further capabilities — no retraining.
Open-set attribution — reject generators never seen in training
Carve out a known-class subset, train on it, then score against the full label map: every class the checkpoint was not trained on is treated as unseen. Detection uses max-softmax-probability under all-patch logit averaging, and is reported threshold-free (AUROC, AU-OSCR).
python scripts/split_known_unknown.py --label-map configs/label_maps/openfake_27class.json \
--n-known 17 --seed 42 --output configs/label_maps/openfake_17known_seed42.json
python train.py --config configs/openfake_17known.yaml --output runs/openset_seed42 --seed 42
python openset.py --checkpoint runs/openset_seed42/best.pt \
--config configs/openfake_27class.yaml --output runs/openset_seed42/evalLineage & unseen-source discovery — unsupervised structure in the features
lineage clusters per-generator mean features hierarchically and reports the
cophenetic correlation; discovery runs UMAP + HDBSCAN over the classes the
checkpoint was never trained on and scores the result against their true
labels (ARI / NMI / purity). Neither uses any lineage or discovery labels.
python cluster.py lineage --checkpoint checkpoints/openfake_27class.pt \
--config configs/openfake_27class.yaml --output runs/openfake_27class/lineage
python cluster.py discovery --checkpoint runs/openset_seed42/best.pt \
--config configs/openfake_27class.yaml --output runs/openset_seed42/discoveryRequires the cluster extra.
Adaptation — admit new generators by fitting a linear head
The backbone is frozen and its penultimate features cached once, so fitting the head takes seconds to minutes instead of a full retrain. All three published regimes share one script; they differ only in what target data feeds the cache.
# full: extend a known-subset backbone to the benchmark's complete class set
python adapt.py --checkpoint runs/openset_seed42/best.pt \
--config configs/openfake_27class.yaml --output runs/adapt_full27 --transplant
# cross-dataset: freeze a backbone trained on one benchmark, fit a head on another
python adapt.py --checkpoint checkpoints/dragon_25class.pt \
--config configs/openfake_27class.yaml --output runs/cross_d2o
# few-shot on GenImage-9 (LIDA's protocol)
python adapt.py --checkpoint checkpoints/openfake_27class.pt \
--config configs/genimage_9class.yaml --output runs/fewshot_10shot \
--shots 10 --train_views 20 --val_split val --test_split val--transplant keeps the trained head row for any class shared by name with
the source checkpoint, so only genuinely new classes start from scratch.
train.py train a classifier (on-the-fly augmentation)
eval.py closed-set evaluation: multi-patch aggregation, robustness, frequency analysis
openset.py open-set attribution (reject unseen generators)
cluster.py lineage recovery + unseen-source clustering
adapt.py adapt a frozen backbone to new classes (full / cross-dataset / few-shot)
datalib/ dataset, patch tiling, augmentation, open-set metrics
models/ the CNN
configs/ one YAML per run, plus label_maps/ ({class_name: folder}, one per benchmark)
scripts/ data download + known/unknown split generator
checkpoints/ download target for the pretrained classifiers
Every entry point takes --config <path> plus its own flags.
@misc{livne2026scalableblackboxmodelattribution,
title = {Scalable Black-Box Model Attribution for Images},
author = {Asaf Livne and Amir Jevnisek and Shai Avidan},
year = {2026},
eprint = {2608.15652},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2608.15652}
}This work builds on the following public benchmarks and baselines:
- DRAGON — 25-generator attribution benchmark.
- OpenFake — 27-generator benchmark.
- GenImage / LIDA — few-shot adaptation protocol.
- AEDR — the white-box eight-model comparison benchmark.
Released under the MIT License.