Created on 2026-05-02 · Last modified on 2026-05-05
The team
Most AI database founders come from the database industry. They apply CRUD intuitions to a category that needs new primitives. I come from automotive software — where systems are not allowed to crash, where ASPICE process maturity and ISO 26262 functional safety are the floor, and where every code review accounts for failure modes before features.
16 years shipping safety-critical software. 7 of those years in Germany leading architecture for safety-critical control units. International ASPICE coaching engagements in UAE and China — coaching distributed teams through complete process rollouts.
Then I started building applications with AI and watched every one stitch 5–8 databases together with duct tape. Every architecture I respected for safety-critical systems — single substrate, atomic batches, deterministic recovery, fuzz validation — none of it applied to the database stack underneath modern AI.
So I built one. Built with automotive-grade rigor — not because it's a marketing line, but because that is what 16 years of muscle memory produces.
Applying automotive-grade engineering rigor to a category — AI databases — that has never seen it before. I have spent more time analysing what happens when systems break than most database founders have spent building them.
Hiring against this round · 5 senior engineers, 1 SDR, 1 AE, 1 marketing lead, 1 ops/CS lead. Team shape built to ship Stage 2 → Stage 3 in 18 months without a senior co-founder gap — Bharath has run distributed teams of similar size and discipline for ASPICE rollouts in Germany, UAE, and China.
The problem
Behind almost every AI product today: rows in Postgres, cache in Redis, search in Elastic, vectors in Pinecone, analytics in ClickHouse, events on Kafka. Six bills, six SDKs, six on-call rotations, six failure modes. The integration cost compounds with every service added — and AI applications are the first generation that needs every shape simultaneously.
Plus six SDKs in the codebase, six dashboards in monitoring, six rotations on-call, and an N²-shaped consistency problem between every pair of stores. The engineering tax compounds long after the bill stops growing.
Why now
AI applications are the first generation needing every data shape simultaneously — rows for state, vectors for recall, reactive subscriptions for live agents, columnar for analytics, NL for the human seam. The 4-service stack doesn't fit anymore.
The solution
A managed-cloud database where every shape — rows, indexes, columnar chunks, vectors, full-text, graph, NL queries — lives at a different hash-keyed prefix on the same engine. One WAL, one recovery, one bill, one SDK. All of these are live today on a single instance.
Verified end-to-end on production EC2 on 2026-05-01: SQL aggregations, HNSW topk with metadata filtering, BM25 ranked search, BFS + weighted Dijkstra all green. Live demo on request.
The wedge · what we eliminate
Today's production RAG splits across 8–10 services. The embedder and the LLM stay yours — those are AI compute, not infrastructure. Everything between them — vector DB, document store, BM25, graph, retrieval cache, retrieval observability — collapses into OriginChain. Seven data-plane services into one engine, one SDK, one bill, one atomic transaction. The cost compression is on the plumbing; the model spend stays where it should.
Embedding model (OpenAI · Cohere · Bedrock · local). OriginChain stores embeddings; it does not generate them. ~$200–$800 / mo for a typical workload.
The LLM itself (GPT-4 · Claude · Llama). OriginChain compiles plans + serves data. Generation is inference, on a separate billing line.
LLM-side observability + eval (LangSmith generation traces · Ragas · Braintrust). EXPLAIN tells you the data path; it doesn't grade the answer.
The honest claim · OriginChain compresses ~60% of a typical RAG bill (data-plane infra · ~$3K → ~$800/mo) while leaving embedder + LLM + LLM-eval as their own line items. Retrieval latency drops 5–10×; end-to-end latency improves by the share of the path OriginChain owns.
Why this isn't another "all-in-one DB" claim · Postgres+pgvector+pg_search lacks graph, NL-in-engine, and reactive primitive. MongoDB+Atlas Vector+Atlas Search same gap. LanceDB / TurboPuffer = vector + BM25, no graph or NL. Vespa = multi-shape but heavyweight, not managed. OriginChain's specific combination — NL compilation in-engine + reactive as a primitive + hash-keyed unification of every retrieval shape, all atomic in one transaction — has not shipped elsewhere.
Multi-tenant architecture
Every key OriginChain stores is hash-prefixed with the tenant ID at the engine level. Two tenants' rows never sit at adjacent keys, can never accidentally collide, and cannot cross-read by accident. Multi-tenancy is a property of the substrate, not a convention you maintain in every query.
Tenant ID is mixed into every key's blake3 hash prefix. Two tenants' keys live in completely different parts of the keyspace by construction. No application-level tenant scoping, no shared-key bugs, no leak class possible.
Each Storm-tier customer gets a dedicated single-tenant EC2 instance. Whisper / Thunder share controlled multi-tenant pools when chosen by the customer. Same substrate, different deployment shapes — choice is per-customer.
Audit log scoped per tenant via the same hash prefix. Per-tenant S3 backup, per-tenant PITR cursor, per-tenant DPA / SOC 2 evidence. The tenant boundary is the unit of compliance, not the cluster.
Why this matters for enterprise sales · multi-tenancy designed in (not bolted on) is the difference between "we have to architect carefully" and "we run customer workloads on isolated keyspaces by construction." Enterprise security reviews accept the second; they spend months on the first.
The vision
Every generation of software gets its default database. The web era picked Postgres + MySQL. Mobile picked DynamoDB + MongoDB. The AI-native generation hasn't picked yet — and the choice is happening now, in the next 18 months. We are building to be the default.
Our target: by 2027, OriginChain becomes the default substrate AI applications reach for. NL-as-a-query-language becomes standard. Hash-keyed unification becomes the architecture pattern engineers teach. The category gets named after the pattern, not after a company — and we aim to own the pattern.
Target: 200 mid-market tenants on Thunder/Storm. 10–20 enterprise contracts post SOC 2. Multi-region GA across 3 regions. Reference customers in trading, agent infra, observability. Series A target: $25 M ARR by month 36 (projected).
7-person team. Public beta open at originchain.ai. 50+ paying tenants. SOC 2 Type 1 complete, Type 2 in motion. First enterprise contracts closed. Series A milestone reached at $3–5 M ARR.
Why OriginChain becomes the default — Postgres won the web era because it was good enough at every shape developers needed and the integration cost was zero. AI apps need a 5–8 service stack today. Whoever ships a substrate that's good enough at every AI-era shape — rows, vectors, full-text, graph, NL, reactive — at zero integration cost wins the same way. All seven shapes are live in our engine today. The architectural decision is made; the only question is who reaches the developer mindshare first.
Market
The global database software market is one of the largest in software, growing double-digits every year. Our beachhead is the segment already paying the poly-store tax — and the next category, AI-native unified, has no entrenched winner.
Traction · engineering proof
Every claim below comes from a drill log, a benchmark, or a CI run — never from a slide writer's imagination. The most rare thing for a pre-seed infra startup is engineering proof. We have it. The next milestone is design-partner conversion, the first paying tenant.
What is honestly aspirational · multi-writer + multi-region + SOC 2 are scoped on the roadmap, not yet shipped. We mark these clearly in code (multi-writer-deferral.md) and we will not market them until they land. We market what we measure.
Business model · unit economics
Every tenant runs on isolated AWS infrastructure sized to their tier. Direct cloud cost is predictable; the AI pass-through metering protects margins on customers whose AI mix runs unexpectedly hot.
| Tier | Price (USD / mo) | Direct AWS cost | Gross margin | Quotas (NL queries · DB calls · egress) |
|---|---|---|---|---|
| Whisper | $99 | ~$9 · t4g.small + 20 GB EBS | ~76 % | 2.5 K · 1 M · 100 GB |
| Thunder | $599 | ~$23 · c7g.large + 100 GB | ~82 % | 25 K · 10 M · 500 GB |
| Storm | $1,499 | ~$85 · 2× c7g + sync replica | ~81 % | 100 K · 50 M · 2 TB |
| Enterprise | from $4,999 | variable · multi-region custom | 70 – 80 % | negotiated |
Plus addons: Vector Search ($79 + $0.0002/topk), SQL Pro ($49), Full-Text Pro ($49), Graph ($59). LTV : CAC target > 4× · net retention target 120 %+ from quota expansion + tier-up.
Go to market · vertical wedges
Whisper $99 · Thunder $599 · Storm $1,499 · Enterprise from $4,999 — every tier is live today. A customer in any of the verticals below can be on the engine within hours, not quarters. What changes by tier is the GTM motion, not product availability. Below is what each industry replaces with OriginChain, the technical buyer in each, and the annual spend pool we compress.
Tick rows + columnar replay + reactive signal fan-out collapse onto one substrate. Removes cross-system joins between hot ticks and analytical backfills.
Vector memory + structured turns + full-text recall + reactive tool-output streams on one engine — removes the four-system glue every agent team rebuilds.
High-volume time-series writes, geospatial keys, reactive alerts, columnar rollups — one substrate replaces the Kafka→Timescale→S3 pipeline.
Fraud similarity + KYC document search + immutable audit logs on one queryable substrate — shorter path from signal detection to compliance evidence.
Logs · metrics · traces · alert subscriptions as different key shapes on one engine — replaces the LGTM sprawl observability teams maintain.
Player state, live leaderboards, retention analytics on one substrate — collapses the Redis-plus-Scylla-plus-warehouse pattern most live-ops teams run.
Records + model embeddings + clinical-note search + HIPAA audit trail — one auditable access path · shorter compliance review.
Catalog + semantic search + behavioural signals + reactive cart on one engine — replaces the Algolia + Pinecone + warehouse merchandising stack.
Sensor history, anomaly embeddings, maintenance alerts on one substrate — modernises the historian + warehouse split industrial AI hits immediately.
Tier ↔ tier mapping · all live today · Beachhead verticals enter via Whisper $99 / Thunder $599 self-serve. Scale verticals upgrade to Thunder / Storm $1,499 with addons (Vector $79, FTS Pro $49, Graph $59). Enterprise verticals close on Storm + Enterprise from $4,999 under custom contract today; SOC 2 Type 1 (in flight) unlocks unrestricted enterprise procurement. Each card above is a workload we compress — not a pipeline we sell into. Every tier ships from day one.
Competition + defensibility
No competitor combines unified substrate + NL-in-engine + production-grade engineering discipline. The substrate is architectural and not retrofittable. The discipline is cultural and compounds with every hire. The market window is open because the AI-native category was created in the last 36 months.
You cannot graft a hash-keyed unified substrate onto Postgres, MongoDB, or ClickHouse. The single-substrate-many-shapes design is the architectural decision; competitors who try to match it ship a rewrite, not a feature. The architectural choice is the moat.
16 years of ASPICE + ISO 26262 rigor applied to an infrastructure substrate. Drill twice, fuzz on every PR, adjust marketing claims to measured numbers. This compounds with every engineer hired and is invisible in competitor product docs. It is also what enterprise DB buyers actually pay for.
Snowflake created cloud-warehouse, MongoDB created document, Confluent created event-streaming — each in a 24–36 month window with no entrenched leader. The AI-native unified category is in that exact window now. Reference customers compound: every name on the wall makes the next sale cheaper.
Why not Postgres + extensions? pgvector ≠ a unified substrate. Why not Snowflake + Cortex? Warehouse, not OLTP. Why not Pinecone? Single shape; still need a primary DB. Why not SingleStore? Heavyweight; no NL primitive; no reactive engine. Each competitor is good at one shape; we are enough good at every shape that a team retires 4–6 services.
How investors make money back
Every previous database to win a category became a $5 – 50 B+ outcome. Below is the actual return math on a $2 M pre-seed at a $10 M post-money cap, fully diluted through Seed → Series A → Series B, against the comparable outcomes that anchor the thesis.
| Comparable | Category won | Outcome | Time | Reference for OriginChain's path |
|---|---|---|---|---|
| Snowflake | Cloud data warehouse | $33 B IPO → $80 B | 8 yrs | Highest exit in the category · top of band |
| MongoDB | Document database | $7 B IPO → $30 B | 10 yrs | Strong enterprise expansion post-IPO · repeatable |
| Confluent | Event streaming | $9 B IPO → $13 B | 7 yrs | Median outcome for category-creating infra · plausible base case |
| Databricks | Data lakehouse | $43 B private | 11 yrs | Mega-private outcome · multiple growth rounds |
| CockroachDB | Distributed SQL | $5 B Series E · 2024 | 9 yrs | Enterprise-only path · no IPO yet |
| Exit scenario | Total exit value | Pre-seed share | Dollar return on $2 M | Multiple |
|---|---|---|---|---|
| Acqui-hire (early exit) | $30 M | 15.80 % | $ 4.74 M | 2.4 × |
| Mid-tier strategic acquisition | $300 M | 15.80 % | $ 47.4 M | 23.7 × |
| Confluent-like IPO (base case) | $9 B | 15.80 % | $ 1.42 B | 711 × |
| MongoDB-like IPO + growth | $30 B | 15.80 % | $ 4.74 B | 2,370 × |
| Snowflake-like outcome (top of band) | $75 B | 15.80 % | $ 11.85 B | 5,925 × |
Bull-case dilution path · Pre-seed 20 % @ $10 M post → after Seed ($20 M @ $250 M post) → 18.4 % → after Series A ($80 M @ $1 B post) → 16.93 % → after Series B ($200 M @ $3 B post) → 15.80 %. The light dilution per round assumes hot up-rounds at each stage — defensible if Whisper/Thunder self-serve traction + Storm enterprise deals compound on schedule. Conservative-case dilution is materially heavier; we present the bull case so the asymmetric upside is visible. Strategic acquirers (downside protection): AWS, Microsoft, Google, Snowflake, MongoDB, Confluent, Databricks. Comparable outcomes are reference frames, not promises.
The ask
~ 18 months of runway with a 7-person team. Funds engineering, sales, marketing, and compliance to take OriginChain from drilled closed beta to public-beta SaaS with 50+ paying tenants, SOC 2 Type 2 in motion, and an active enterprise pipeline.
$2 M SAFE @ $10 M post-money cap. Closing Q3 2026. Min check $50 K · max single check $750 K. Seeking 1–2 lead investors with infra-DB experience.
Intros to trading-desk + AI-infra teams for design-partner round. Hiring help on senior systems engineers. EU enterprise relationships post SOC 2.
Email info@originchain.ai for a 30-min walkthrough + live demo. Technical product deck + drill postmortems + DECISIONS log open under NDA at pitch.originchain.ai.
The window is ~ 18 months. After that, AI teams have chosen their substrate and switching costs lock in. We are ready to ship. We need allocation.
© 2026 Silicoyn Technologies Pvt Ltd · OriginChain is a managed-cloud database product · Pre-seed investor pitch · Confidential