Open this publication in new window or tab >>Show others...
2026 (English)In: Machine Learning with Applications, E-ISSN 2666-8270, Vol. 23, article id 100806Article in journal (Refereed) Published
Abstract [en]
Detecting advertisements in digitized newspapers is a key step in large-scale media analytics and digital archiving. However, variations in layout, typography, and advertisement design across publishers and time periods cause significant domain shifts that reduce the generalization ability of supervised detectors. This paper presents AdAPT, a confidence-guided pseudo-labeling pipeline for unsupervised domain adaptation in advertisement detection. The proposed method leverages both advertisement-free (Null) and advertisement-containing pages from unlabeled target domains to generate reliable pseudo-labels. By retraining a YOLO-based detector using labeled source data combined with filtered pseudo-labeled target samples, AdAPT achieves robust adaptation without requiring manual annotation. Experiments conducted on two unseen newspapers (Adresseavisen and iTromsø) demonstrate that Null-based pseudo-labeling provides the most stable and accurate adaptation, yielding up to 38% error reduction compared to the baseline. The results highlight AdAPT as a simple, scalable, and annotation-efficient solution for maintaining high-performance advertisement detection across diverse newspaper collections.
Place, publisher, year, edition, pages
Elsevier BV, 2026
Keywords
Cross-domain advertisement detection, Deep learning, Domain adaptation, Object detection, Pseudo labeling
National Category
Computer Sciences
Identifiers
urn:nbn:se:miun:diva-56539 (URN)10.1016/j.mlwa.2025.100806 (DOI)2-s2.0-105027856969 (Scopus ID)
2026-02-032026-02-032026-02-03