What is an AI Tech Stack? A Founder's Guide to the Layers

A prototype becomes a product when a team can make its best demo work reliably for every customer. For AI products, the components that make that moment repeatable live in everything around the model, which is what an artificial intelligence (AI) tech stack describes.

Early architecture choices in that surrounding system are the expensive ones to reverse. This guide covers the five layers of the stack, the choices that shape each one and what the whole thing costs to run.

What is an AI Tech Stack

An AI tech stack is the set of layers that turn a model into a working product. Compute, data, the large language model (LLM), the orchestration code your application runs on and the evaluation layer that scores output make up those layers. Monitoring runs across all of them. Most of what decides whether the product works is outside the model itself.

Noisy retrieval or a bloated prompt can produce a bad answer long before it reaches the model. A routing error can do the same, and the model is the component you'll swap most often. Your data pipeline, test sets, caching and model-independent guardrails last longer.

The work the AI does for the customer drives which layers get built and which get rented. A team extracting fields from insurance claims needs a different data layer than a team shipping a coding agent, even when both call the same hosted model. Support deflection leans on retrieval quality; coding agents lean on orchestration and tool control. Those differences become clearer when you separate the stack into its operating layers.

The 5 Layers of an AI Tech Stack

Nearly every production AI product has five layers, whether the team names them or not. Each has a hosted default a two-person team can rent and a self-managed option that usually arrives later. The tools below are common at seed and Series A.

1. Infrastructure and Compute

At seed stage, compute usually means an application programming interface (API) bill. Most seed stage teams call a managed model API and run application code on a container service such as Google Cloud Run or a serverless provider such as Modal. Kubernetes can wait until an infrastructure team exists to operate it. Engineers spend their time on the product instead of on cluster maintenance, and heavy traffic is what makes rented graphics processing unit (GPU) economics worth a second look.

2. Data and Retrieval

In retrieval-augmented generation (RAG) products, output quality depends more on clean, well-chunked source data than on which database stores the vectors. The pipeline is the same everywhere: load documents, split them into chunks, generate embeddings and store them for similarity search. Labeling and curation are the slow, hands-on part of that work. CRV-backed Encord consolidates annotation across video, audio, text and sensor data, with dataset curation in one place. Teams that treat this layer as a one-time import let the index go stale.

3. Models and Inference

Hosted APIs from Anthropic, Google and OpenAI are where nearly every seed stage product starts, and the price spread inside a single provider is the first cost lever you'll pull. Lower-cost models can handle routine, high-volume work for much less than flagship reasoning models. A practical cost pattern is to route classification and extraction to the cheap tier and reserve the flagship for reasoning. Open weights such as DeepSeek V3.2 or Llama 4 that you host yourself come later in a company's life.

4. Orchestration and Application Logic

Orchestration is the code that decides what context the model sees and what happens with the answer, and most of your product lives there. Engineering teams fix workflow sequences in advance, while agents let the model choose the next tool or action as the task unfolds. Architecture should enforce human approval gates; LangGraph's interrupt primitive pauses a run at a node until someone approves, which a prompt line can't guarantee. CRV-backed Vercel covers the product surface and deployment, and its software development kit (SDK), AI SDK 5, added stopWhen and prepareStep for agent loop control. Vercel's gateway also fronts a broad model catalog with automatic fallback.

5. Evaluation and Observability

Most stack diagrams bury evaluation, and most founders skip it until a regression reaches production. Tracing tools such as Langfuse, LangSmith and Weights and Biases Weave record what happened on each call. You can inspect inputs, outputs, cost, token use and latency to see where a bad answer came from. The other half is a golden dataset of reviewed inputs and expected outputs, small enough to run on every pull request, with a rubric or an LLM judge scoring it. Drift is why this can't wait: accuracy of generative pre-trained transformer 4 (GPT-4) on a prime-number task fell from 84 percent to 51 percent between March and June 2023.

AI Tech Stack vs. Traditional Software Stack

The architecture diagram looks familiar and the operating assumptions underneath it do not. Habits you built shipping software as a service (SaaS) break in a few specific places once a model sits in the request path:

  • Cost behavior: Every request carries variable compute cost, so cost of goods sold (COGS) scales with usage as you add customers. Your heaviest customers are also the most expensive to serve.
  • Output behavior: The same input can return different output, so correctness becomes a scored range. Your test suite scores output against that range.
  • Data dependency: Data the team keeps current sets the ceiling on quality, and code alone can't raise it.
  • Release process: Evaluation suites replace unit tests as the gate before shipping, because exact-match assertions fail on valid responses and pass on invalid ones. A prompt change runs against the golden dataset and doesn't merge if scores regress.
  • Generative AI versus classical machine learning (ML): The older ML stack meant custom-trained models and hand-built feature sets; in the generative stack, prompts and retrieval configure a foundation model, with full retraining as a last resort. Behavior you used to fix in code now moves into data and prompts, with tests governing both. That shift changes what has to be true before you ship.

How to Choose Your AI Tech Stack

Every layer offers the same trade: rent a managed version now, or spend engineering time to control it yourself. At seed stage, teams should nearly always choose the option that keeps engineers on the product, and each choice below comes with the condition that flips it.

Hosted Model APIs vs. Self-Hosted Models

Teams should use hosted models until data residency or volume economics force the change. Against a managed open-weight provider, self-hosting usually reaches break-even only at sustained high volume, and the engineers who'd babysit that GPU are the ones who'd otherwise ship product. Residency is increasingly something you can buy from some providers, though regional coverage, eligible models and pricing remain uneven. The remaining exceptions are proprietary fine-tuned weights no third-party API will run and regulated workloads in healthcare, finance or legal that require prompts to stay on premises. Outside those two cases, hosted stays the default.

Retrieval vs. Fine-Tuning

Retrieval often comes first, because it is a cheaper option to try when facts change frequently and can be updated by refreshing the data source rather than retraining the model. A new document lands in a retrieval index in minutes, while fine-tuning takes hours to days and assumes a data scientist you haven't hired. Fine-tuning fits cases where knowledge is stable and you need consistency at volume. Consistency of that kind usually means format, tone and terminology, and even then fine-tuning pairs with retrieval.

Frameworks vs. Direct API Calls

Plain application code around a direct API call is the faster path for a single-model product with a handful of tool calls. Orchestration frameworks such as LangGraph earn their complexity when you need durable execution and persistence across long agent sessions. Human-in-the-loop interrupts you'd otherwise build from scratch provide another reason to use one. You give up visibility in exchange, since framework behavior can hide behind class parameters and changing it can require editing the framework's own source.

Managed Services vs. Self-Managed Infrastructure

A managed vector database or model gateway buys back engineering weeks, the scarcest input a seed stage team has. Running pgvector on the Postgres instance you already have costs nothing extra at small scale, while a dedicated service such as Pinecone adds a separate usage-based bill. An OpenAI-compatible gateway such as LiteLLM or a unified client library such as the AI SDK turns a model change into a configuration edit. Everything else works fine as a managed service at this stage.

What an AI Tech Stack Costs to Run

Inference is part of COGS, which pulls AI gross margins into the 50 to 60 percent band against the 80 to 90 percent SaaS teams underwrote. The same compression shows up in public company filings. C3.ai's generally accepted accounting principles (GAAP) gross margin declined from 59 percent in fiscal third quarter 2025 to 32 percent in fiscal first quarter 2027. Cost per query is the unit economics metric worth tracking from the first customer, because inference and compute can outweigh payroll at AI-native companies.

Seat counts are a poor proxy for inference load: one customer may make a few short calls while another runs long, multi-step workflows all day. That mismatch is pushing AI coding products toward usage-based pricing as providers pass inference costs to their most active customers. Inference also keeps getting cheaper, and the price to query a model at GPT-3.5-level performance fell from $20 per million tokens in November 2022 to $0.07 by October 2024. Our inference pricing guide covers how much a single provider announcement can cut a model's price.

The AI Tech Stack Mistakes That Stall Startups

Team size doesn't change which stack decisions go wrong, only what they cost to unwind later. These recurring failures come from making stack decisions in the wrong order:

  • Provisioning for scale you haven't earned: Reserved GPU commitments and a Kubernetes cluster before steady traffic buy idle hardware and an operations burden a three-person team can't carry.
  • Shipping with no evaluation layer: Quality regressions surface as support tickets instead of failed test runs. Without production-shaped test cases, a model update can score well before release and still reach users with behavior severe enough to force a rollback.
  • Coupling application logic to one provider's format: Prompts you tuned for one model's quirks don't transfer, and every call that uses a vendor's client library turns a model swap into a rewrite.
  • Measuring per-query cost after the bill arrives: Retries and multi-step agent chains consume tokens a prototype budget never modeled, and staging environments add to that use. Teams find out what a query costs a month after they spent the money.
  • Treating the data layer as plumbing: Ingestion pipelines assembled piecemeal create undocumented dependencies that make failures hard to trace. Cost per query and the evaluation set both belong ahead of any tool decision. Product requirements should drive vendor choices, and they rarely do when tools come first.

Cost per query and the evaluation set both belong ahead of any tool decision. Each is cheap to put in place early and expensive to retrofit once the stack is built.

What Separates an AI Tech Stack That Ships From One That Stalls

The stacks that hold up at early stage AI companies are narrow. Teams build them around a single job, own the data and evaluation layers and rent everything else behind one thin abstraction. A founder who can state cost per query and name the eval set that gates each release is describing a business. They should also be able to explain which layer they'd swap first.

A small, deliberate stack is not a sign of underinvestment at seed and Series A. More often it means the team decided what the product had to do before deciding what to buy. We back founders who assume the model will be replaced and put their engineering into the layers that outlast it.

If you're an early stage founder looking for a partner who will get into the weeds on how you build your data pipeline and eval set, reach out to us to see if we'd be a good fit.

Frequently Asked Questions About the AI Tech Stack

What is the difference between an AI stack and an AI tech stack?

In practice the two terms describe the same thing: the layers of infrastructure, data, models, orchestration and evaluation that turn a model into a product. Which term someone uses tends to reflect who is talking; the stack itself stays the same.

Do you need a vector database in your AI tech stack?

At seed stage, Postgres with pgvector usually meets most teams' needs. For those workloads, it keeps documents and embeddings in a single write and costs nothing beyond the database instance you already pay for. A dedicated service such as Pinecone or Qdrant earns its keep when scale, write intensity or latency requirements exceed what a general-purpose database can meet.

Does an AI tech stack require machine learning operations engineers?

At seed stage, software engineers with an interest in ML can share the machine learning operations (MLOps) work, so most teams defer a dedicated MLOps hire. The same person often wires the model API, writes the retrieval pipeline and sets up tracing. A dedicated hire makes sense once ML runs in production with real users and someone needs to own that system full time.

What does a first AI tech stack look like at seed stage?

A first stack usually holds one frontier model API (Anthropic or OpenAI) and a vector-enabled Postgres database for retrieval. Plain application code or a light framework such as LangChain or LlamaIndex connects them, with deployment on a FastAPI service or Vercel. Langfuse covers tracing and evals, or LangSmith if you're already on LangChain or LangGraph. The smallest available tier of each service is the right starting point.

Congrats to Lotus AI and Outtake on Making Forbes Next Billion-Dollar Startups List

CRV proudly co-led Lotus AI’s Series A and our firm led Outtake’s Series A and joined the board in February 2025. We also backed Outtake during its Series B, so we’re thrilled to see both teams make this year’s list.” to “CRV proudly co-led Lotus AI’s Series A and our firm led Outtake’s Series A, joined the board and backed Outtake during its Series B, so we’re thrilled to see both teams make this year’s list.

CRV invests in founding teams at the beginning of their journeys, leading Seed and Series A rounds in amazing companies. We’ve backed more than 750 companies early on including DoorDash (another Next-Billion alum), Mercury and Vercel.

Congrats to both Lotus AI and Outtake on being named to Forbes’ Next-Billion Dollar Startups list.

Cookie Preferences

Your Privacy Matters to Us

We use cookies and similar technologies on this site, employed by CRV and our partners, to support core features and help us understand how visitors engage with our content. For details, please review our Privacy Policy