A

Interpretable Visual Feature Discovery using Multiple Instance Learning: A Case Study of Colonial Korean Print

Zenodo (CERN European Organization for Nuclear Research)

Abstract

Reproducibility bundle for the article Interpretable Visual Feature Discovery using Multiple Instance Learning: A Case Study of Colonial Korean Print, forthcoming in Computational Humanities Research (Cambridge University Press, ISSN 2977-8158). Cite the article Van de Pol, A., Prokic, J., & Mol, A. (forthcoming). Interpretable Visual Feature Discovery using Multiple Instance Learning: A Case Study of Colonial Korean Print. Computational Humanities Research, X(X). Cambridge University Press. Volume, issue, and page range will be added once the article is in print. The deposit contains: Source code archive (code-v1.0.0.tar.gz) with training, evaluation, and interpretability pipelines. Trained AttriMIL Flash model weights for both 16x16 and 8x8 patch configurations (5-fold each, optimizer state stripped). Trained ConvNeXtV2 Base and Swin S3 Base 224 baseline weights under both held-out-book GroupKFold and page-random KFold protocols (5-fold each). Page-level metadata for the 57,583-page corpus, with canonical printshop labels and source catalog links. Fold definitions for both evaluation protocols, including the page-to-book mappings used for the GroupKFold splits. Per-fold evaluation results for every released model. Interpretability artifacts: 8x8 patch embeddings with attention weights, UMAP and t-SNE coordinates, HDBSCAN cluster assignments, and a compact cluster-assignment artifact. Documentation: dataset card, manifest with SHA-256 checksums, and a deposit-level README. The page images themselves are not redistributed and remain governed by the Hyundam Mun'go collection terms; per-page URLs are recorded in the dataset metadata. Licensing: The source code archive is released under the MIT License. All other artifacts (dataset metadata, fold definitions, model weights, evaluation results, interpretability artifacts) are released under CC-BY-4.0.

Authors 3

  1. Leiden University

    Affiliation as printed

    Leiden University

  2. Leiden University

    Affiliation as printed

    Leiden University

  3. Leiden University

    Affiliation as printed

    Leiden University

Cited by 0 stored of 0

No patents citing this paper on Lens.org (checked 2026-10-11).

References 0