Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Mythos Metadata Backup Engine

Executive Summary

The Mythos Metadata Backup Engine is a headless system service engineered to facilitate the secure extraction, transformation, and archiving of multi-tenant enterprise layout and configuration metadata. Operating autonomously, the system executes complex data lifecycle management routines without requiring direct user intervention or synchronous client polling.

At its core, the application is a high-throughput, concurrent, and fault-tolerant batch migration pipeline built upon the robust foundations of Spring Boot 3.x and Spring Batch 5.x. The primary objective is to reliably stream high-volume relational configuration states, serialize them into a portable JSON-Lines format, compress the artifacts, and securely archive the resulting payloads to AWS S3 storage infrastructure. The architecture is explicitly designed to handle rigorous data processing constraints while maintaining strict consistency and isolation protocols.

System Architecture & Data Flow

The following diagram illustrates the end-to-end data processing lifecycle of the Mythos Metadata Backup Engine, detailing the orchestration from initial database extraction through chunk processing, fault isolation, and final S3 archival.

flowchart TD
    %% Define Styles
    classDef boundary fill:#f3f4f6,stroke:#9ca3af,stroke-width:2px,stroke-dasharray: 5 5;
    classDef container fill:#e1f5fe,stroke:#0288d1,stroke-width:2px,color:#01579b;
    classDef volume fill:#fff3e0,stroke:#f57c00,stroke-width:2px,color:#e65100;
    
    subgraph DockerHost ["Docker Compose Ecosystem / Host Machine"]
        
        %% Persistent Volumes
        PgVolume[("postgres_data<br/>(Persistent ACID Storage)")]:::volume
        LsVolume[("localstack_data<br/>(Mocked S3 State)")]:::volume
        TmpVolume[("Local Host Backup Dir<br/>(/tmp/mythos-backups/)")]:::volume
        
        subgraph NetworkBridge ["Mythos Virtual Bridge Network"]
            Engine["Mythos Engine Container<br/>(Spring Boot 3.x / Java 17)"]:::container
            DB[("PostgreSQL 15 Container<br/>(Relational State)")]:::container
            S3{{"LocalStack AWS Container<br/>(S3 Storage Emulator)"}}:::container
        end
        
        %% Mappings and Data Flow
        Engine <==>|JDBC / Port 5432| DB
        Engine ==>|AWS SDK / Port 4566| S3
        
        DB -.->|Mounts| PgVolume
        S3 -.->|Mounts| LsVolume
        Engine -.->|Bind Mounts| TmpVolume
    end
    
    style DockerHost fill:#ffffff,stroke:#374151,stroke-width:3px;
    style NetworkBridge fill:#f8fafc,stroke:#94a3b8,stroke-width:2px,stroke-dasharray: 5 5;
Loading

Detailed Component Interaction Flow

To complement the macro system architecture above, the following sequence illustrates the micro-level internal orchestration running within the Mythos Engine Container during a scheduled trigger:

sequenceDiagram
    participant Trigger as Job Launcher
    participant L1 as BackupJobLifecycleListener
    participant DB as PostgreSQL (JPA)
    participant Reader as RepositoryItemReader
    participant Processor as MetadataBackupProcessor
    participant Writer as BackupArchiveWriter
    participant S3 as AWS S3 (LocalStack)

    Trigger->>L1: Trigger metadataBackupJob
    L1->>DB: INSERT BackupJob (Status: PROCESSING)
    
    rect rgb(235, 245, 255)
        note right of Reader: Multi-Threaded Chunk Processing (Core: 5)
        Reader->>DB: SELECT * FROM metadata_records WHERE backed_up=false LIMIT 100
        DB-->>Reader: Return Page Chunk
        
        Reader->>Processor: Yield items
        Processor->>Processor: Validate Payload & Compress
        
        alt Validation Failed
            Processor--x DB: INSERT BackupAuditLog (SkipPolicy)
        else Validation Passed
            Processor->>Writer: Forward Valid Records
            Writer->>Writer: Serialize to JSON-Lines
            Writer->>LocalDisk: Write backup-chunk-XXX.jsonl
        end
    end
    
    Trigger->>L1: afterJob() Invoked
    L1->>LocalDisk: Compress JSON-Lines into ZIP Archive
    L1->>S3: PUT / Multipart Upload Archive
    S3-->>L1: Return 200 OK (S3 URI)
    L1->>DB: UPDATE BackupJob (Status: COMPLETED, URI)
Loading

Core Architectural Capabilities

The following explicit system design paradigms are implemented within the engine to ensure deterministic and scalable execution:

  • Chunk-Oriented Processing: The pipeline leverages a stream-lined extraction model using a transactional repository pattern with strict page isolation. Data is processed in configured block sizes (chunks) to optimize heap memory consumption and minimize network payload overhead.
  • Multi-Threaded Step Scaling: The system features concurrent performance optimizations using an explicit ThreadPoolTaskExecutor executing parallel background processing. This vertical scaling allows partitioned threads to simultaneously process independent data chunks, drastically reducing total execution time for large datasets.
  • Fault Tolerance and Resiliency: To prevent total pipeline collapse during isolated data anomalies, the engine utilizes programmatic exception handling using localized skip logic. This is coupled with real-time database auditing via custom step listeners, ensuring that malformed payloads are safely bypassed and recorded for operations analysis.
  • Data Integrity Guardrails: The architecture guarantees transport-layer validation by combining SHA-256 state signatures mapped at the record level with Content-MD5 header verification during network transit to prevent bit-rot and ensure cryptographic consistency.

System Topology and Technology Stack

The technology footprint of the Mythos engine is selected to prioritize non-blocking I/O, enterprise maturity, and seamless cloud integration.

Component Implementation Framework Engineering Rationale
Batch Mechanics Spring Batch 5.x Provides standardized chunk-oriented processing, integrated retry/skip policies, and transaction boundary management.
Data Persistence Hibernate and Spring Data JPA Offers high-level ORM abstractions over relational state tracking, minimizing raw SQL maintenance.
Network I/O AWS Java SDK v2 (Netty NIO) Utilizes a non-blocking asynchronous HTTP client (Netty) for high-concurrency S3 streaming and multipart uploads.
Relational Tracking PostgreSQL (or H2 in-memory) Chosen for ACID compliance and advanced indexing capabilities required for tracking massive audit trails.
Application Runtime Spring Boot 3.x (Java 17) Ensures a modular, auto-configured application lifecycle suitable for modern container orchestration environments.

Database Schema and State Management

The internal state of the application is maintained across three primary relational tables designed for optimal read-write performance during batch cycles.

  • metadata_records: Represents the source of truth for the multi-tenant configuration artifacts. It contains localized fields such as tenant identifiers, metadata classifications, and the raw layout payloads.
  • backup_jobs: Serves as the authoritative tracking registry for execution lifecycles. It records the temporal state of the batch execution (e.g., pending, processing, completed, failed) alongside the final S3 bucket URL mapping.
  • backup_audit_logs: A diagnostic persistence layer utilized by the skip listeners. It stores granular item-level failures, capturing the exact step phase, the record identifier, and the underlying exception message for subsequent telemetry.

To optimize connection I/O during high-volume sequential extraction, the schema assumes the use of a PostgreSQL partial index on the backed_up column (WHERE backed_up = FALSE). This architectural choice dramatically reduces the indexing footprint and accelerates the RepositoryItemReader query execution plans when scanning for unprocessed configuration layouts.

Local Development and Cloud Emulation Setup

The ecosystem is fully containerized to ensure identical parity between local development and remote deployment targets. The architecture abstracts cloud physical layers using a containerized LocalStack image (pinned to version 3.8.0) to emulate AWS S3 endpoints seamlessly without incurring live cloud resource costs or requiring active network credentials.

To run the environment locally, ensure the Docker daemon is active and execute the following commands.

  1. Initialize the isolated cluster network and build the application containers:
docker compose up --build -d
  1. Provision the required LocalStack S3 bucket prior to triggering the batch execution:
docker compose exec localstack-aws awslocal s3 mb s3://mythos-backups
  1. Query the relational audit logs using the PostgreSQL container utility to verify fault-tolerance triggers:
docker compose exec postgres-db psql -U mythos_user -d mythos -c "SELECT step_name, record_identifier, error_message FROM backup_audit_logs;"
  1. Execute validation queries via the AWS CLI directed at the emulator to confirm the final artifact upload:
aws --endpoint-url=http://localhost:4566 s3 ls s3://mythos-backups/backups/ --recursive

Code Quality, Testing, and Extension

The codebase is fortified by a comprehensive verification framework, specifically utilizing S3 integration testing suites backed by JUnit 5 and Testcontainers. This ensures that custom storage interactions, such as Content-MD5 hashing and multipart boundaries, are rigorously evaluated against realistic mock infrastructures.

Furthermore, the integration logic implements a dual-credential strategy. While local configurations explicitly pass mocked keys for the LocalStack emulator, the production profiles omit static secrets entirely. By relying on the AWS DefaultCredentialsProvider chain, the application automatically resolves temporary security tokens. This renders the compiled JAR instantly deployable and secure within cloud environments, such as AWS ECS or EKS, via standard IAM role associations.

About

A headless metadata backup service for enterprise data processing

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages