TensorX
返回文献探索

Paper · arXiv 2608.00440

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis

Zhishan Zou

10 upvotesAugust 1, 2026arXiv 预印本
AI 摘要

Poplar is a reproducible pipeline that synthesizes and curates human-centric image-text datasets through structured specification, realism-adapted rendering, and vision-language inspection.

image generatorvision-language reviewrealism-adapted image generatorstructured vision-language review

Abstract

Recent image generators can synthesize convincing human-centric images, yet producing a useful collection remains different from producing a single successful image. A human-centric dataset must cover varied people and contexts, avoid implausible attribute combinations, preserve an everyday photographic character, and expose quality-control decisions at scale. We present Poplar, a reproducible Specify--Render--Inspect pipeline for human-centric image dataset synthesis. Specify samples structured attributes under commonsense constraints and verbalizes them as photography-oriented prompts. Render uses a realism-adapted image generator across composition-aware aspect ratios and retries obvious technical failures. Inspect applies a single structured vision--language review to each candidate, preserving the original prompt while rejecting intrinsic image defects or material prompt mismatches. Using Poplar, we construct Poplar-9K: 9,401 curated human-centric image--text pairs retained from 11,765 reviewed candidates (79.9\% acceptance). We release the dataset together with the pipeline, configurations, immutable generation prompts, and auditable inspection records as a compact resource for building customizable human-centric collections.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号