Architecture
Architecture & Core Engine
LotusGen is engineered for low memory consumption and high throughput, decoupling metadata configuration from chunked data generation pipelines.
Technical Stack
- Frontend: React 19, TypeScript, Vite, Tailwind CSS v4, Phosphor Icons.
- Backend API: FastAPI with SQLite metadata storage for projects and schema definitions.
- Data Engine: Polars and PyArrow chunked execution pipeline for bounded RAM usage during multi-million row exports.
- Expression Sandbox: Custom Python AST parser validating and executing user formulas in a restricted namespace.
- Dependency Graph: Directed Acyclic Graph (DAG) validator ensuring topological ordering of column evaluations and table dependencies.
Processing Pipeline
- Schema Validation: The system parses the table schema and constructs a dependency graph across columns and referenced tables.
- Topological Sort: Columns are sorted topologically so computed columns only reference previously generated columns.
- Chunked Streaming: Data generation runs in configurable chunk batches. Batches are streamed directly to disk or network sockets via PyArrow writers without buffering entire datasets in memory.
- Seed Propagation: A root project seed deterministically initializes random state generators across all worker routines.