Skip to content

Architecture

Architecture & Core Engine

LotusGen is engineered for low memory consumption and high throughput, decoupling metadata configuration from chunked data generation pipelines.


Technical Stack

  • Frontend: React 19, TypeScript, Vite, Tailwind CSS v4, Phosphor Icons.
  • Backend API: FastAPI with SQLite metadata storage for projects and schema definitions.
  • Data Engine: Polars and PyArrow chunked execution pipeline for bounded RAM usage during multi-million row exports.
  • Expression Sandbox: Custom Python AST parser validating and executing user formulas in a restricted namespace.
  • Dependency Graph: Directed Acyclic Graph (DAG) validator ensuring topological ordering of column evaluations and table dependencies.

Processing Pipeline

  1. Schema Validation: The system parses the table schema and constructs a dependency graph across columns and referenced tables.
  2. Topological Sort: Columns are sorted topologically so computed columns only reference previously generated columns.
  3. Chunked Streaming: Data generation runs in configurable chunk batches. Batches are streamed directly to disk or network sockets via PyArrow writers without buffering entire datasets in memory.
  4. Seed Propagation: A root project seed deterministically initializes random state generators across all worker routines.