AI engineering and infrastructure at Nuqta
What's under the surface makes the difference.
Agents, multi-model architecture and inference infrastructure we run ourselves.
187 technologies across 17 areas
- In our products
- Evaluated
- R&D
- Planned
Models & LLMs
Multi-model by design: no single-model lock-in; each task goes to the right model.
Mibyan 4.1- OpenAI models & APIs
- DeepSeek
- Qwen 3.x · Qwen3.5 / 3.6 35B-A3B
- Kimi K2.6Evaluated
- Gemma 12B / 31BEvaluated
- Hermes Agent
- Open-source LLMs
- Long-context (128K+)
- Model routing & fallback
- Multi-provider architecture
- Bring your own model / provider
Agentic AI
Agents that act: tools, memory, sub-agents and continuous evaluation.
- Multi-agent systems
- Tool / function calling
- Structured outputs
- Sub-agents
- AI memory
- Evals
- Research agents
- Browser agents
- Computer-use agents
- Workflow automation
- AI coding workflows
- AI-generated documents & presentations
- Voice agentsR&D
Knowledge & retrieval
Answers grounded in the organization's own sources, not guesses.
- RAG
- Knowledge bases
- Vector & semantic search
- Embeddings
- Reranking
- Document & file understanding
- Long-context retrieval
- Enterprise knowledge retrieval
- Conversation memory
- Persistent agent context
- Source-grounded responses
- Data analysis with AI
Model serving & inference
We serve models ourselves and measure them: speed, concurrency and cost.
- vLLM
- OpenRouter
- OpenAI-compatible APIs
- Self-hosted serving
- Streaming inference
- High-concurrency inference
- Quantization · FP8 · NVFP4
- KV / context optimizationEvaluated
- Throughput, latency & load benchmarks
- Cost / quality routing
Mibyan Inference Gateway
GPU & private AI infrastructure
From cloud GPUs to on-premise and sovereign AI.
- NVIDIA H100
- NVIDIA H200
- NVIDIA L40S
- NVIDIA RTX 4090
- NVIDIA RTX PRO 6000
- NVIDIA A100 80GBEvaluated
- NVIDIA B200Planned
- Local & cloud GPU inference
- Multi-GPU architectureEvaluated
- Dedicated GPU servers
- Private & sovereign AI
- On-premise / private cloud / hybrid
Cloud & platforms
We deploy where it suits the customer, including technical discussions on Oman-hosted inference.
- RunPod
- Supabase Cloud
- Coolify
- Vercel
- Railway
- Hostinger
- Linux servers
- Containerized workloads
- National cloud (Oman-hosted inference)R&D
- High availability & scalable inference
- CI/CD
Backend
Streaming APIs, background workers and an architecture built to grow.
- Node.js 24 LTS
- TypeScript
- Fastify 5
- Python 3.13
- FastAPI
- REST & streaming APIs (SSE)
- OpenAI-compatible API design
- Webhooks
- API gateways
- Modular monolith
- Event-driven workflows
- Sandboxed execution
- Zod 4
- OpenAPI 3.1
Frontend & web
Arabic-first, bilingual interfaces — from chat to editors and dashboards.
- React 19
- Next.js 16 (App Router)
- Vite
- TanStack Query
- Zustand
- React Hook Form
- Tailwind CSS v4
- HeroUI v3
- next-intl
- TipTap
- Shiki · react-markdown · GFM
- RTL & Arabic-first UI
- PWA
- Real-time & AI chat interfaces
- Artifact / editor interfaces
AI frameworks & MCP
MCP is central to how we integrate: agents connect to tools and systems with clear permissions.
- Vercel AI SDK
- @openrouter/ai-sdk-provider
- Agent orchestration
- Custom model routers
- MCP servers & clients
- Database MCP
- Developer-tool MCP
- External-service MCP
- Authentication-aware MCP
Data layer
Multi-tenant databases with row-level isolation and full audit trails.
- PostgreSQL
- Supabase (Auth · Storage · Realtime)
- Row-Level Security
- pgvector-style search
- pgcrypto
- pgmq
- Kysely
- Prisma
- Redis
- BullMQ
- Multi-tenant data models
- Usage metering
- Audit logs
Identity & security
Role-based access, protected secrets, sandboxed execution and data residency where needed.
- OAuth 2.0 · Google OAuth
- SSO / LDAP integration
- RBAC & organization roles
- API keys & secrets management
- Private storage
- Rootless sandbox workers
- Enterprise data isolation
- Data residency architecture
- Audit logging
- VAPT-ready delivery
Desktop & mobile
Agents that run on the machine itself: files, terminal and browser.
- Electron · Electron Forge
- macOS apps (DMG, notarization)
- Local file & terminal access
- Browser automation
- Computer use
- Local agent execution
- Capacitor
- PWA
- OAuth via system browser
Engineering workflow & DevOps
Agent-assisted development and safe delivery across separate environments.
- Git · GitHub (protected main)
- Codex
- Claude Code
- Cursor
- AI coding agents
- Docker (multi-stage)
- Rootless containers
- KubernetesEvaluated
- Coolify
- Staging & production environments
- Health checks
Quality & observability
We test product and model together, and watch performance and cost in production.
- Vitest
- Playwright
- React Testing Library
- Testcontainers
- End-to-end & API testing
- GPU / model benchmarks
- Concurrent-request tests
- OpenTelemetry
- Pino (structured logs)
- Token, usage & cost tracking
- Inference metrics
Document generation
From conversation to finished file: Word, PowerPoint, Excel and PDF, with AI editing.
- python-docx
- python-pptx
- openpyxl
- LibreOffice (headless)
- Playwright PDF rendering
- HTML/CSS documents
- Inline AI editing
- Preview / editor systems
Communication & automation
WhatsApp, email, calendar and CRM, with escalation to a person when needed.
- WhatsApp automation
- AI customer-service agents
- Lead qualification & follow-up
- Email & calendar automation
- CRM automation
- Human-in-the-loop
- Multi-channel architecture
- Twilio · phone numbersR&D
- Real-time voice AIR&D
Computer vision & edge
Cameras and on-site processing, in projects like ParkEye and SAHL.
- Camera-based vision
- Parking-space detection
- Real-time occupancy monitoring
- Smart attendance
- Edge / on-site processingR&D
- Mac mini M4 (SAM runtime)
- Local servers & GPUs
Designed to integrate with
Systems our products are built to connect to — not a list of deployments.
- SAP
- Oracle
- Microsoft Dynamics
- Odoo
- Microsoft 365
- SharePoint
- Google Workspace
- ERP
- CRM
- HR
- Email & calendar
- Databases
- Internal & open APIs
- Procurement systems
- Business portals
- Call centers
Where this stack runs
- Mibyan Chat
- Mibyan 4.1
- Mibyan API
- Mibyan Desktop
- SAM
- Mibyan Enterprise
- Mibyan Tender
- Mibyan PWA
- WhatsApp AI
- SAHL
- ParkEye
- Healthcare CRM
- Education AI
- Enterprise CRM
Benchmark
~111–157tokens/sec
Measured in internal tests on specific models and configurations. Not a general figure for all models.
Deployment options
- CloudFully managed by Nuqta.
- Private cloudInside your cloud environment.
- On-premiseOn your own infrastructure.
- APIBuild your products on Mibyan.
- Bring your own modelChoose the provider or model.
Mibyan Platform
An OpenAI-compatible API: change the base URL, keep your code.
# OpenAI-compatible
from openai import OpenAI
client = OpenAI(base_url=MIBYAN_BASE_URL, api_key=MIBYAN_API_KEY)
client.chat.completions.create(model="mibyan-4.1", messages=[...])