AI stack
Our AI technology stack: models, frameworks and deployment
The models, frameworks and infrastructure we use to build chatbots, autonomous agents and knowledge assistants — integrated with the same secure web and backend stack as the rest of your software.
Technologies in our AI stack: OpenAI GPT, Claude, Gemini, Llama, DeepSeek, Mistral, LangChain, LlamaIndex, CrewAI, AutoGen, MCP, Ollama, Pinecone, Weaviate, Python, FastAPI, Node.js, Docker, Kubernetes, Redis.
The AI technology stack
Mainstream tools, chosen per use case
We are model-agnostic by design. Each category below is a toolbox we know deeply, so we can pick what fits your data, budget and privacy requirements.
AI models
We choose per use case — accuracy, speed, cost and data residency — and design the application so the model can be swapped as better options appear. Open-weight models can run entirely on your own servers.
- OpenAI GPT
- Claude
- Gemini
- Llama· open-weight
- DeepSeek· open-weight
- Mistral· open-weight
Frameworks & protocols
Orchestration for retrieval, multi-step agents and tool use, plus local model serving.
- LangChain
- LlamaIndex
- CrewAI
- AutoGen
- MCP
- Ollama
Vector databases
Semantic search over your documents and records, the foundation of accurate retrieval-augmented generation.
- Pinecone
- Weaviate
AI backend
Python services for model and data work, connected to our Node.js and NestJS application backend.
- Python
- FastAPI
- Node.js
DevOps
Containerised, reproducible deployments that scale with demand and queue long-running AI work.
- Docker
- Kubernetes
- Redis
Integration
AI that plugs into our standard web and backend stack
AI is never a separate island. Every AI feature sits behind the same NestJS API, with the same validation, permissions, logging and tests as the rest of your software — so it is secure and maintainable by any developer in our pool.
- Step 01
Your apps & systems
Web app, mobile app, CRM, ERP or internal tools — the places your people already work.
- Step 02
NestJS API · /v1
Authentication, role and record-level permissions, input validation and rate limits on every AI request.
- Step 03
AI service · Python + FastAPI
Retrieval, prompts, agents and tool calls, with long-running work queued through Redis and BullMQ.
- Step 04
Models & vector store
Cloud model APIs or local models via Ollama, grounded in your data through a vector database.
Deployment options
Cloud, on-premise or hybrid — your choice
Where your AI runs is a business decision about privacy, cost and capability. We support all three and help you choose.
Cloud
Hosted model APIs and managed infrastructure for the fastest start and access to the most capable models.
Best for: Fast pilots, variable workloads and tasks that need frontier-model quality.
On-premise
Open-weight models served on your own hardware or private cloud, so prompts and documents never leave your network.
Best for: Strict confidentiality, data-residency rules and predictable high-volume usage.
Hybrid
Sensitive data stays on local models while less sensitive or harder tasks route to cloud models, under one policy.
Best for: Organisations balancing privacy, cost and capability across many use cases.
FAQ
AI technology stack: frequently asked questions
Have an AI use case in mind?
Book a free 30-minute consultation. We will assess feasibility, suggest the right models and deployment option, and outline a measurable pilot.
Keep exploring
Related pages
- AIAI ServicesLearn more
- AILocal LLM DeploymentPrivate models where your data never leaves your servers.Learn more
- AIAI Agent DevelopmentAutonomous, auditable agents that complete real work.Learn more
- AIAI Knowledge BaseAnswers grounded in your documents with RAG.Learn more
- TechnologyWeb & Backend StackNext.js, NestJS, TypeScript, PostgreSQL and Prisma.Learn more