Skip to content

Commit a661fbe

Browse files
committed
Merge remote-tracking branch 'upstream/main' into dsaa2026
2 parents 0433765 + c7f42e7 commit a661fbe

3 files changed

Lines changed: 63 additions & 0 deletions
Lines changed: 42 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,42 @@
1+
---
2+
layout: post
3+
title: "Leveraging Synthetically Generated Data for Real Estate Document Classification"
4+
date: 2026-04-29
5+
author: Tobias Deußer
6+
categories: [Research]
7+
description: false
8+
---
9+
10+
This is a short summary of our paper **"Leveraging Synthetically Generated Data for Real Estate Document Classification"** by *Tobias Deußer, Gregor Ramien, Nico Weber, Maximilian Meidinger, Max Hahnbück, Christian Bauckhage, and Rafet Sifa*, published in the proceedings of the 2025 IEEE International Conference on Big Data. See [here](https://ieeexplore.ieee.org/abstract/document/11400789) for the published version and [here](https://bonndoc.ulb.uni-bonn.de/xmlui/handle/20.500.11811/13972) for the open-access bonndoc version.
11+
12+
Document classification in regulated industries like real estate and banking is hindered by a fundamental tension: the models that could automate workflows require large amounts of labeled training data, yet the sensitive nature of the documents makes collecting and annotating real examples difficult or legally fraught. In this work, we address this challenge by building a pipeline that generates realistic synthetic training data without ever exposing real documents.
13+
14+
Our focus is on two document types central to German mortgage financing workflows: the *Nachweis Kindergeld* (Child Support Certificate) and the *Individueller Sanierungsfahrplan* (Refurbishment Roadmap). Rather than collecting and anonymizing real examples, we construct Jinja-based templates derived from manual analysis of genuine documents and populate them with a combination of rule-based (Faker library) and LLM-generated content. A third catch-all class, *Sonstige* ("Other"), is generated entirely by prompting an LLM with a diverse set of everyday document types to serve as a realistic negative class.
15+
16+
The pipeline proceeds as follows:
17+
18+
![Figure 1: The complete data synthesizing and training pipeline for our document classifier.](/assets/blog/leveraging-synthetically-generated-data-for-real-estate-document-classification-training_setup.jpg)
19+
*Figure 1: The complete data synthesizing and training pipeline for our document classifier.*
20+
21+
1. **Template Design:** Real documents are analyzed manually to extract layout patterns, static text, and variable fields.
22+
2. **Placeholder Filling:** Simple fields (names, dates, addresses) are filled via the Faker library; semantically rich fields are filled by an LLM to preserve document coherence.
23+
3. **Augmentation:** Documents can be reordered, rephrased, or degraded with synthetic OCR noise to increase variation.
24+
4. **Classifier Training:** A BERT-based encoder with a linear classification head is fine-tuned on the resulting synthetic pages.
25+
26+
We evaluate the trained classifier on a manually annotated real-world test set of 52 documents (240 pages), compiled by domain experts, the model never saw any real documents during training. The composition of this test set is shown below:
27+
28+
| Category | # Pages | # Documents |
29+
|---|---|---|
30+
| Empty Page | 4 | 2 |
31+
| Child Support Certificate | 53 | 32 |
32+
| Refurbishment Roadmap | 120 | 12 |
33+
| Other | 63 | 6 |
34+
| **Total** | **240** | **52** |
35+
36+
**Key Results:**
37+
38+
- **Basic 3-class setup:** Page-wise average precision of **88.3%**, demonstrating strong zero-shot-to-real-world transfer.
39+
- **Fine-grained 15-class setup** (each page template as a distinct class): Page-wise AP of **81.4%**, showing the approach scales to more granular classification.
40+
- **Reduced template scenario** (75% of Refurbishment Roadmap templates): AP of **80.8%**, confirming that competitive performance can be achieved even with incomplete template coverage, thus reducing manual effort.
41+
42+
The main limitation is that template creation remains manual and time-consuming, and synthetic data may not fully capture the long-tail variability of real-world documents. Future work will target fully automated template generation using structure-aware LLMs, extension to layout-aware models, and broadening the approach to information extraction tasks beyond classification.
126 KB
Loading

‎teaching/mmdii26/index.md‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -15,6 +15,27 @@ This course, offered as part of the *Master's Program in Human Centered Intellig
1515
* Analyze media data to derive actionable insights and support decision-making in real-world domains such as digital marketing, financial data analysis, and text/document analytics.
1616
* Tackle key challenges in media analytics, including ethical issues, model interpretability, and efficient use of computational resources.
1717

18+
19+
| Lecture | Date | Lecture Content |
20+
| -----------:| -----------| --------------------------------------------------------------------------------------|
21+
| Lecture 1 | 30.4.2026 | Course Intro |
22+
| Lecture 2 | 7.5.2026 | Data Mining Applications 1 |
23+
| Lecture 3 | 21.5.2026 | Data Mining Applications 2 |
24+
| Lecture 4 | 11.6.2026 | Representation Learning with Transformers 1: Intro and Training |
25+
| Lecture 5 | 18.6.2026 | Representation Learning with Transformers 2: Textual Embeddings |
26+
| Lecture 6 | 18.6.2026 | Representation Learning with Transformers 3: Text mining and NLP |
27+
| Lecture 7 | 25.6.2026 | Representation Learning with Transformers 4: Multimodality |
28+
| Lecture 8 | 25.6.2026 | Representation Learning with Transformers 5: Big Data Engineering in the age of LLMs |
29+
| Lecture 9 | 9.7.2026 | RL and LLMs 1 |
30+
| Lecture 10 | 16.7.2026 | RL and LLMs 2 |
31+
| Lecture 11 | 23.7.2026 | Information Extraction and Practical Applications |
32+
33+
34+
## Lecure Notes
35+
- Lecture 01: [ZIP Download](https://github.com/AppliedMachineLearning-Lab/mining_media_data_II_SS26/raw/refs/heads/main/MMD_II_SS26_Lecture_01_CourseIntro.zip)
36+
37+
38+
1839
## Lecturers
1940

2041
- [Rafet Sifa](/rafetsifa/)

0 commit comments

Comments
 (0)