vibcreates Los Angeles, CA · 34.05°N

Vibhav Gupta — founder & data engineer

I build systems of record for data that disagrees with itself.

Ten years across data platforms and reliability engineering — seven of them as founder & CEO of CannMenus, the pricing-and-inventory record for North American cannabis retail. This year I pointed the same method at three stranger domains: satellite catalogs, exoplanet archives, and the tennis court.

Open to — senior / staff data-platform & product roles Taking on — fractional data engineering & AI consulting
Proof, not promises querying live systems…
tracked objects resolved · orbital
orbital element sets · orbital
exoplanet candidates · exodossier
source assertions kept · exodossier
The work

Four domains, one method, all real.

01
Retail

CannMenus

Founder & CEO · 2019 — present

Cannabis retail had thousands of dispensary menus and no ground truth. CannMenus scrapes, normalizes, and reconciles them into SKU-level pricing and inventory intelligence covering roughly 90% of licensed dispensaries in the U.S. and Canada, refreshed hourly. Bootstrapped to real revenue, raised angel capital, then acquired and integrated a Canadian competitor's data business. Led a team of four; ran the stack end to end — Python, Kafka, TimescaleDB, BigQuery, Kubernetes, Airflow.

~90%
of licensed dispensaries covered
hourly
inventory + price refresh
1
acquisition led & integrated
7 yrs
operating, founder-led
02
Orbit

Orbital Economy Intelligence

Built solo · 2026 · live & refreshed nightly

Space-Track, CelesTrak, GCAT, and the UCS registry openly disagree about who owns what in orbit — and nearly everyone resolves it by quietly picking one source. OEI keeps all four: an identity graph with per-attribute provenance, a replayable merge log, and the disagreements surfaced as the product. 99.3% of tracked objects carry at least one cross-source conflict; here, that's a feature with receipts.

70,080
satellites, one resolved identity each
9.7M
orbital element sets, hypertabled
99.9%
operator attribution coverage
4,400+
live conflicts surfaced, not averaged
03
Deep sky

ExoDossier

Built solo · 2026 · live & refreshed nightly

The exoplanet archives disagree on radius, temperature, even whether a candidate is a planet at all — flips that change a world's habitability. ExoDossier reconciles the NASA Exoplanet Archive, ExoFOP, KOI, and Gaia into one provenance graph and generates cited vetting dossiers per candidate. It also ships the first MCP server for exoplanet data, so AI agents can query the sky directly.

20,341
host stars cross-matched
23,033
candidates tracked
684k
per-attribute source assertions
7,900+
radius / disposition / Teff conflicts
04
Baseline

SplitStep

Inventing solo · 2026 · in field testing, v0.17

Tennis training from phones you already own — no sensors, no subscription, no cloud in the sensing path. Two unsynchronized phones agree on a ball impact through NTP-style clock sync and time-difference-of-arrival; the pairing itself filters the noise. A 1.16 MB ball detector I trained from scratch exports one set of weights to Android, iOS, and the research bench. Sound knows when, vision knows where. Provisional patent drafted.

2
phones, zero extra hardware
1.16 MB
ball-tracking CNN, trained from scratch
sub-ms
cross-device clock agreement
1
provisional patent drafted
Method

The same move, every time.

iFind the disagreement

Where authoritative sources contradict each other is exactly where a system of record is worth building. Most pipelines average the conflict away; I catalogue it — it's the most honest signal in the data.

iiKeep the receipts

Every value carries its source, its timestamp, and the precedence rule that made it win. Merges are logged and replayable. Trust is an audit trail, not a vibe.

iiiGrade against ground truth

OEI and ExoDossier ship gold-standard evaluations with published accuracy. For SplitStep I re-benchmarked a published 83% prior-art claim and measured 31% — then built something better and graded it the same hard way.

ivShip it live

Working software over decks. Everything above runs in production, refreshes on a schedule, and answers real queries — on infrastructure I operate myself.

Before this

A decade of data platforms and reliability.

2023University of Chicago — Genomic Data Commons · Senior SWE — test automation and cloud orchestration for national cancer-research data
2020 – 22SpotHero · SWE — E2E frameworks, Pact contract testing, load & capacity work with SRE
2019 – 20DAIS Technology (acq. Origami Risk) · Senior SWE — CI/CD test frameworks across concurrent products
2018 – 19Rewards Network · SDET — QA architecture for a streaming-data re-platform
2017 – 18Uptake · SWE — CI/CD and automation for a streaming big-data platform
2015 – 17Networked Insights (acq. AmFam) · QA Automation — re-architected regression, 5× faster
eduDePaul University · B.A. Economics, Honors Program
Contact

Two ways to work with me.

Need a system of record —
or the person who builds them?

vibhavgupta2@gmail.com github.com/vxg4120 linkedin.com/in/vib

This site is hand-written static HTML on a $7 server in Helsinki that I administer myself. The live numbers above are real queries against my databases, stamped at request time. View source — there's nothing to hide.