LotusGen
LotusGen
LotusGen is a high-throughput, low-memory dummy and synthetic data generation platform designed for software engineers, data analysts, and QA teams. It enables the creation of complex, relational, and reproducible tabular datasets at scale in CSV and Apache Parquet formats.
Key Capabilities
- High-Performance Streaming Pipeline: Built on Polars and PyArrow chunked data processing, keeping memory consumption strictly bounded regardless of total row count.
- Relational Integrity: Inter-table and intra-table DAG validation ensures correct topological generation order and prevents circular dependencies.
- Secure Custom Logic: Custom AST (Abstract Syntax Tree) expression sandbox allows powerful dynamic column formulas without unsafe code execution.
- Strict Reproducibility: Master integer seed control ensures every random draw, UUIDv7 sequence, and categorical sampling is deterministic.
- Dual Export Formats: Direct export to compressed CSV and Snappy-compressed Apache Parquet.
Documentation Sections
- Architecture & Core Engine: Internal design, streaming core, and safety sandboxing.
- Table Archetypes & Data Model: Record, Aggregate, Reference, and Collection table structures.
- Column Types & Roles: Column roles and supported statistical data types.
- Formula Expressions: Syntax, operators, math functions, and conditional exclusions.
- Output & Determinism: Seed management, chunked streaming, and file formats.
- Deployment & Self-Hosting: Docker Compose, environment variables, and local setup.
- REST API Reference: Endpoints for automated generation pipelines and integration testing.