Session 20, “Digital Archaeological Collections as AI Training Data”, brought together researchers working at the intersection of archaeology, digital collections, and artificial intelligence. Chaired by MAiA members Vera Moitinho de Almeida, Nevio Dubbini, Aurore Mathys, and Gabriele Gattiglia, the session featured 11 presentations and lively discussions on the opportunities and challenges of using archaeological data to train AI systems.
The session focused on questions that are central to the MAiA COST Action. What makes a high-quality AI training dataset in archaeology? How do archaeological recording standards align with the technical requirements of machine learning? How can researchers address the biases embedded in archaeological archives, digitisation strategies, and historical documentation practices? And how can AI applications remain transparent, reproducible, and ethically accountable?
Across the presentations, speakers explored how digital archaeological collections can support new forms of analysis while also highlighting the risks of reproducing existing gaps, interpretations, and biases in the archaeological record. Discussions emphasised the importance of FAIR and well-documented datasets, interdisciplinary collaboration, and critical reflection on the methods used to develop AI applications for archaeology.
Several presentations showcased emerging workflows and technical approaches for preparing archaeological data for AI-driven analysis. Topics included pottery reconstruction, segmentation of wall paintings and thin sections, legacy data digitisation, 3D scanning workflows, deep learning applications for faunal remains, and generative modelling for archaeological materials.
The session also highlighted the growing importance of open and reusable digital comparative collections. High-quality training data remains one of the major challenges for AI in archaeology, and contributors stressed the need for shared standards, sustainable infrastructures, and collaborative approaches across institutions and disciplines.
The presentations included:[b]
- “The MAiA project and Digital Comparative Collections as AI Training Data for Archaeology” — Vera Moitinho de Almeida, Nevio Dubbini, Aurore Mathys, Gabriele Gattiglia
- “Experimental AI Applications for Rapid Archaeological Legacy Data Distribution” — John Wallrodt
- “Natural Schema Evolution vs Machine Learning Readiness: A Case Study from the Stone-Masters Project” — Maciej Krawczyk
- “From Legacy Data to Training Data: AI-Driven and Open Archaeological Workflows: the example of pottery” — Lorenzo Cardarelli, Julian Bogdani
- “Automated Segmentation and Integration of Avifaunal Bone Image Datasets Using Deep Learning-Based Mask Generation” — Nevio Dubbini, Gabriele Gattiglia, Beatrice Demarchi, Lisa Yeomans, Marco Pavia, Paola Sansone, Ramazan Parmaksız, Ayşe Ataş Hooglugt
- “From Microscopy to AI-assisted Petrography: Preparing Archaeological Thin Sections for Segmentation in TagLab” — Elisabetta di Virgilio, Giorgio Gosti, Diego Ronchi, Marco Callieri
- “Documenting Fragmentary Wall Paintings through AI-Based Segmentation in TagLab. Toward a Standardised Workflow and FAIR Archaeological Datasets” — Caterina Paola Venditti, Diego Ronchi, Silvia Fortunati, Giorgio Gosti, Marco Callieri
- “Harnessing AI to unlock legacy data: the AutArch experience and beyond” — Maxime Brami, Kevin Klein, Felix Riede
- “Generative Modeling for Potteries: A GMM-GAN Based Framework for Completion and Clustering of Pottery Fragments” — Suhui Liu
- “From Pots to Points and Back Again: A 3D Scanning and Generative Machine Learning Workflow for Pottery Studies” — Dries Daems, Jitte Waagen, Mason Scholte, Mikko Kriek
- “Carian pottery geographical differentiation using meta-learning, transfer learning, and signal processing based neural network hybrid architectures” — Deniz Kayikci, Juan Antonio Barceló
