DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

1State Key Laboratory of Content Cognition, People's Daily Online 2Nanyang Technological University 3Tongji University 4Nanjing University of Science and Technology 5National University of Singapore 6University of Science and Technology of China

Abstract

Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and localized manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and localized object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. In addition to the benchmark, we propose FPD (Fake-Preferences Detector), a diagnostic-driven strong baseline that encourages forgery-sensitive representation learning. FPD combines a dual-pathway architecture with a cascaded compression classifier to progressively extract and compact forgery-relevant cues, and further introduces a fake-preference sampler to reduce excessive bias toward the real class. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91–96% balanced accuracy on GenImage drop to 60–76% on FakeBench and 54–66% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and localized manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods.

Dataset Sample Examples

Dataset Statistics

Dataset sample visualization

Experimental Results

FakeBench Results

Methods SD3.5 FLUX.1 FLUX.2 Z-Image Qwen-Image N.B. 2 GPT-I. 2 Avg.
R.AccF.Acc R.AccF.Acc R.AccF.Acc R.AccF.Acc R.AccF.Acc R.AccF.Acc R.AccF.Acc
UniFD 98.902.70 98.720.19 98.830.50 98.801.03 98.630.51 99.590.00 98.540.00 49.78
LOTA 99.830.13 99.820.02 99.890.35 99.810.11 99.860.03 100.000.00 99.6498.91 57.03
DRCT 85.1629.05 85.0421.78 84.7227.65 85.2641.82 85.1730.12 82.4522.86 86.5022.26 56.42
FatFormer 99.642.12 99.550.04 99.540.22 99.590.94 99.540.49 100.000.00 99.6412.04 50.95
NPR 61.8562.66 61.5979.09 60.4066.24 61.5573.95 60.9480.88 63.6759.18 62.0491.24 67.52
SAFE 94.942.48 95.023.15 95.232.41 94.681.88 94.913.88 97.145.31 94.1677.74 54.49
AIDE 94.0013.29 94.2920.40 93.8812.15 94.1621.95 94.2222.63 95.925.71 94.8915.33 55.34
Effort 38.7470.04 39.7474.31 38.3069.16 39.4582.49 39.0988.76 40.0085.31 40.5197.08 60.21
ForgeLens 99.960.19 99.930.66 99.900.31 99.920.23 99.917.37 100.000.00 100.0098.91 57.66
LTD 99.772.64 99.741.36 99.772.48 99.725.06 99.774.88 100.000.41 100.0068.98 56.04
Aligned 99.8149.30 99.8116.27 99.843.01 99.8612.12 99.8521.39 100.0042.04 99.2710.58 60.94
B-Free 89.2255.74 88.968.49 90.015.35 89.0637.28 89.392.78 91.020.82 89.050.73 52.70
RPOBE 87.6796.65 87.8519.18 88.4628.66 87.5871.62 88.1377.43 89.8042.45 86.8636.13 70.60
OmniAID 88.8749.04 88.6937.20 88.9516.40 88.1239.57 88.6945.43 89.8354.87 89.0529.19 63.84
DDA 86.1976.34 85.8130.97 86.9222.70 85.9963.12 86.4456.47 86.1245.31 84.3151.09 70.27
DINO-Det 94.1173.86 93.9686.82 94.2451.48 93.8857.26 94.2358.92 94.6929.39 93.8054.38 76.50
FPD (Ours) 83.3484.63 87.6789.78 89.6163.21 88.7068.39 83.6871.90 94.6929.39 90.6253.76 76.78

Table 1. Accuracy (%) comparison on FakeBench. We report Fake Accuracy (F.Acc), Real Accuracy (R.Acc), and Average Accuracy (Avg.) to evaluate detector performance and generalization capability. N.B. 2 and GPT-I. 2 denote Nano Banana 2 and GPT-Image 2, respectively.

ManipulationBench Results

Methods FLUX-Fill Random FLUX-Fill Object FLUX.2-klein Qwen-Edit Step-Edit N.B. 2 GPT-I. 2 Avg.
UniFD50.16 / 50.1639.02 / 50.8338.43 / 50.7239.61 / 50.3939.10 / 50.7237.64 / 50.2637.45 / 50.2340.20 / 50.47
LOTA49.92 / 49.9237.79 / 49.9237.31 / 49.9238.88 / 49.9737.84 / 49.8937.32 / 50.00100.00 / 100.0048.43 / 57.08
DRCT51.00 / 51.0042.51 / 50.8641.60 / 50.5644.79 / 52.1741.31 / 50.1349.92 / 57.2157.82 / 63.9946.99 / 53.70
FatFormer49.86 / 49.8637.90 / 49.9537.56 / 50.1139.06 / 50.0538.36 / 50.2637.32 / 50.0045.91 / 56.9640.85 / 51.02
NPR45.09 / 45.0940.99 / 46.0244.98 / 49.8152.40 / 55.1744.00 / 48.7752.70 / 55.5381.92 / 80.3351.72 / 54.38
SAFE52.22 / 52.2241.63 / 51.9439.00 / 49.9938.99 / 48.9637.40 / 48.2637.97 / 49.4684.94 / 86.6747.45 / 55.35
AIDE51.13 / 51.1339.86 / 50.6538.16 / 49.6350.45 / 58.6348.31 / 57.5443.54 / 54.4351.65 / 60.4346.15 / 54.63
Effort45.60 / 45.6046.81 / 45.7052.62 / 50.6759.98 / 56.9856.05 / 53.5264.16 / 60.8575.90 / 68.9957.30 / 54.61
ForgeLens50.01 / 50.0137.85 / 49.9837.46 / 50.0541.25 / 51.8939.41 / 51.1741.41 / 53.1791.82 / 93.4948.45 / 57.10
LTD49.88 / 49.8837.76 / 49.8837.37 / 49.9739.00 / 50.0437.95 / 49.9837.48 / 50.0466.57 / 73.4043.71 / 53.31
Aligned50.51 / 50.5138.46 / 50.4437.94 / 50.4267.27 / 73.1838.51 / 50.4355.16 / 64.1445.91 / 56.9647.68 / 56.58
B-Free62.36 / 62.3657.67 / 69.9757.12 / 63.5777.92 / 79.9558.59 / 64.6134.70 / 45.4342.47 / 51.8655.83 / 62.53
RPOBE61.67 / 61.6756.58 / 62.9359.58 / 65.4692.10 / 91.5155.12 / 61.8365.47 / 69.6180.34 / 82.4767.26 / 70.78
OmniAID52.66 / 52.6644.34 / 53.3647.75 / 56.4088.10 / 88.5054.57 / 61.4667.32 / 71.8256.55 / 63.0858.75 / 63.89
DDA54.54 / 54.5446.52 / 54.8255.51 / 62.2989.44 / 89.2056.77 / 62.8856.63 / 62.4880.20 / 81.9662.80 / 66.88
DINO-Det52.22 / 52.2242.41 / 52.9743.35 / 53.9962.95 / 69.0145.69 / 55.5340.43 / 52.0465.14 / 71.5550.31 / 58.18
FPD (Ours)52.56 / 52.5646.01 / 54.0949.36 / 57.8279.02 / 80.4754.17 / 60.7046.25 / 53.2273.86 / 77.1657.32 / 62.29

Table 2. Average / Balanced accuracy (%) comparison on ManipulationBench. FLUX-Fill Random and FLUX-Fill Object denote FLUX-Fill using random and object masks, respectively. Qwen-Edit and Step-Edit denote Qwen-Image-Edit and Step1X-Edit-v1p2, respectively.

For more detailed results, please see the paper .

BibTeX

@article{dailybench,
  title={DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models},
  author={Xin Jiang, Hao Tang, Junyao Gao, Meiqi Cao, Fei Shen, Dongming Zhang, Yongdong Zhang},
  journal={arXiv preprint arXiv:2607.24016},
  year={2026}
}