A graph-native Anti-Money Laundering (AML) benchmark dataset.
Dubai-focused synthetic persona generation using DAG-based probabilistic sampling — fork of MatrAIx (arXiv:2608.04205). 48 dimensions across residency, employment, housing, culture, digital behavior, and health. 1M personas calibrated against DSC, DHA, RTA, KHDA, and GMI data sources.
Generate a realistic mock API straight from your TypeScript types - zero config, one command.
Synthetic Persona Pretraining (SPP): Alignment from Token Zero — umbrella repo for spp-data, spp-training, spp-evals
A synthetic customer book measured two ways: 91.6% or 80.8% three-year retention depending on when the fields are read
API REST gratuita con datos falsos en espanol para prototipos, demos y frontends.
A high-fidelity synthetic retail POS transaction dataset containing 2,000,000+ rows in SQL, JSON, CSV, and Prolog formats. Perfect for database stress-testing and machine learning.
Generates original ARC-AGI-1-style tasks distribution-matched to the public eval set.
The MNIST of gas sensing — open-source simulation engine + ML benchmark for optical spectroscopy. 10 molecules, 9 tasks, 12+ baselines.
Cinematic 3D airport operations simulation: trigger disruptions, watch 600 synthetic passengers react, and deploy a deterministic optimizer.
Interactive educational reproduction of the Coldcard weak-RNG seed attack: vulnerable PRNG -> BIP39/BIP32 -> batch-match funded addresses -> recover mnemonics. 100% synthetic data, education & authorized testing only.
High-fidelity, bare-metal industrial simulation environment written in Nim to generate Data Matrix (ECC200) datasets with real-world physical defects for YOLO-OBB training
MerchantBench is a 365-day, order-level benchmark for evaluating the long-term coherence of LLM agents in seller-side e-commerce operations.
🗓️ The hardest life-admin benchmark for agents — lawsuits, escrow shortfalls, apartment hunts, exams. 20 long-horizon tasks × 20–30 stages across 10 domains and 21 services, scored by 1247 atomic checks that read backend state, not prose. Bilingual zh/en. Current 20 tasks are relatively easy.
Synthetic data generation, post-training, and E2B benchmark evaluation infrastructure.
This package implements a deterministic, closed, cross-vendor experiment for eliciting, classifying, preserving, verifying, and summarizing Semantic Void observations.
TUI tool for synthetic dataset generation and LLM distillation.