Embeddings + Search Modernization
Last verified: June 2026· engagement
We replace or augment legacy search (Elasticsearch, Algolia) with semantic search and RAG — often starting as a small AI feature and growing into a retrieval platform.
- Outcome
- A new semantic search / retrieval layer in production
- Timeline
- 6–12 weeks
- Pricing
- $40–120k build + $4–10k/mo ops
- Buyer
- VP Eng at content/doc-heavy companies

The problem
Your users can't find anything. Legacy search (Elasticsearch, Algolia, basic SQL LIKE) misses intent, and the answers are often wrong. Meanwhile, you have embeddings on the roadmap but no clear path from "we want semantic search" to a production retrieval system.
What we do
- Stand up a retrieval layer over your existing data — vector store, hybrid lexical + semantic, the right chunking strategy.
- Ship RAG that cites its sources, so users can trust the answers and your team can debug failures.
- Wire evals into the retrieval loop so quality regressions are caught before they reach users.
- Roll out incrementally — one feature or one surface at a time, with metrics that show the win.
How it fits together
What you get
Built on our open source
ormai — Gives AI agents database access without the risk — scoped, governed access for agents.
Common questions about this engagement
Do you work under NDA?
Yes — we sign your mutual NDA before any data or repo access. For audits we prefer read-only access to start; for builds we work in a clean repo under your ownership.
Will we own what you ship?
Always. You own the code, the runbooks, the dashboards. We are explicitly set up to hand off and transition out, not to create dependency.
Vendor-agnostic — what does that mean in practice?
We integrate with what you already run — OpenAI, Anthropic, open-weight models on your cloud, your CI/CD, your observability stack. If a hosted vendor solves it, we will not reinvent it; if a self-hosted tool is the right answer, we will not pretend the hosted one is.
How is this different from a Big Four consulting deck?
We are the engineers doing the work, not analysts handing recommendations to a different team. The deliverable is working software in your repo, not a slide deck.
Can you work with our in-house AI team instead of replacing them?
Yes — most of our engagements pair with an internal owner and ramp them up to run the system after we leave. Many of our best engagements start with "we hired an AI team, help us get them productive."
Guides for this work
What is RAG (retrieval-augmented generation)?
RAG retrieves the passages relevant to a question and puts them in the prompt, so the model answers from your data with citations instead of relying on what it happened to memorise during training.
LLM evalsHow to build LLM evals
Evals turn 'the model feels better' into a number you can gate a deploy on — they are the test suite that makes every prompt, model, and pipeline change safe to ship.
Let’s scope it on a call
Thirty minutes with an engineer. We’ll tell you straight whether this is the right first move for your team.