Our services · UAE & GCC

LLM on-Premise

All-in-One Sovereign Agentic AI Platform

Language models and AI agents on your own servers. They read mail, work through contracts and tenders, keep the CRM and close tasks by morning — while company data never leaves your walls. We install, tailor to your processes, train people and support.

0 bytes leave the perimeter Unlimited: no per-request or per-seat fees Arabic · English · Russian Prices published — the only ones in the segment
data leaving the perimeter
0
bytes
first agent in production
10–14
weeks
hardware configurations
33
from $4k
per-request fee
∞
requests
01
Why you need this

Public AI is fine until it meets your data, your regulators and your budget

Companies in the UAE and the region handle data that cannot be sent abroad, and volumes at which the cloud API bill becomes a budget line. A local platform solves both — and adds a third thing: agents that do the work rather than just answer questions.

77%

Staff are already sending your data to public AI

77% of GenAI users paste company data into it, 22% paste personal and payment data, 82% do it through personal accounts with no control at all. Bans do not work — your own AI that is pleasant to use does.

Source: LayerX, 2025
$7.2M

The average breach in the Middle East

SAR 27M per incident; shadow AI adds another $670k to each. 97% of those hit had no AI access controls.

Source: IBM Cost of a Data Breach, 2025
AED 1M+

Data export is already penalised

Health data stays in the UAE (law 2/2019, up to AED 1M). Bank customer data only with CBUAE approval. ADGM up to $28M, Saudi PDPL up to SAR 5M, doubled on repeat.

Source: Chambers, 2026
20 000 000

Chats OpenAI was ordered to hand over

A US court in 2025 ordered all ChatGPT logs preserved and 20M chats handed to plaintiffs. Enterprise contracts were carved out — the rest were not. There is nowhere to send a local model an order.

Source: terms.law, 2025
88%

Of UAE executives: switching AI vendors is hard

84% call a week of vendor downtime critical. Vendors retire models (GPT-4.5 after 4.5 months) and double prices. An open model on your server will not vanish or get dearer.

Source: IBM IBV, 2026
$611

A month per employee — what active teams spend on cloud AI

The median company spends $11 per person — it does not need its own server, and we will say so. Active teams spend $611, teams running agents up to $7,450. With 50 active staff the Business server pays back within a quarter. The calculator below runs on your numbers — honestly, including the cases where the API is cheaper.

Source: Ramp AI Index, 2026
853

Roles covered by one AI agent at Klarna

$60M saved a year, two thirds of contacts, 82% faster replies. An agent that does the work — not a chat that answers questions.

Source: Klarna Q3 2025
74%

Of UAE companies plan to run AI locally

72% fear data leaking into third-party AI tools. Public cloud as the main inference environment fell from 56% to 41% in a year.

Source: Dell, 2026 · Broadcom, 2026
02
What each role gets

One working day with the platform

Not a feature list, but what the agents do from the first month. Every action waits for approval where a person should decide.

07:30
CEO

The brief before you walk in

  • The numbers from your systems, the one off its line, and three decisions for today
  • Any company document — ask in chat, the answer cites the page
09:00
Finance

Receivables with the policy applied

  • Who is overdue, what the credit policy requires, reminder letters drafted
  • Invoices checked against contracts and VAT before filing; month-end pack by the first
11:00
Sales & CRM

The lead showed up in the CRM by itself

  • From a mail or a call — a card, a score, the next step and a draft reply
  • The quarter forecast from the real pipeline, not from gut feel
14:00
Operations

A complaint arrives — the ticket is already open

  • A workflow built from blocks read the mail against the policy, opened the ticket, drafted the reply
  • A person presses send. Nothing leaves without a person
16:00
Legal & Tenders

80 pages read before the day ends

  • What they ask, deadlines, knock-out clauses — pages cited
  • A first draft of the response from your own documents; a contract checked against your policy
18:00
HR

CVs ranked, onboarding assembled

  • Candidates ranked against the role; staff questions on policy answered in their language
  • A newcomer gets a plan, access and materials assembled automatically
02:13
Risk & Audit

Everything the AI did tonight — on record

  • Every action in a tamper-evident log: visible if a single character changed
  • What the agent is unsure about waits for a person on the phone — with context
scroll sideways
03
How it is built

Seven layers — all inside your perimeter

One server or a cluster. No layer calls outside: models, the document index, agents and the log live on your hardware. Tools connect over MCP — the open protocol handed to the Linux Foundation in December 2025.

Your perimeter
L1People & channelshow people use it
Open WebUI / LibreChatTelegram · WhatsAppMicrosoft 365Google WorkspaceMobile appVoice
L2Agents & workflowswho does the work
Named agents per roleLangGraphn8n / DifyTask boardSchedules & triggersHuman approval step
L3Tools over MCPwhat they connect to
MCP (Linux Foundation)Jira · ConfluenceSalesforce · HubSpot1С · Bitrix24Mail, calendar, telephonyYour CRM
L4Company knowledge (RAG)what the agents know
Docling · MarkItDownPDF · Word · Excel · PPT · OCRQwen3-Embedding · BGE-M3pgvector → Qdrant → MilvusAnswers cite the source
L5Models & inferencewhat they think with
Qwen3 · Qwen3.5DeepSeek-V3.2Llama 4Gemma 4Mistral 3gpt-ossFalcon-H1-Arabic (TII)Jais 2 (G42)vLLM · SGLangLiteLLM
L6Trust & governancewho may do what
SSO · RBAC (Keycloak, AD)NeMo Guardrails · Llama GuardPII masking (Presidio)Tamper-evident logSecrets vault
L7Hardware & operationswhere it all lives
GPU server in your server roomKubernetes / DockerPostgres · S3Langfuse · OpenTelemetryEvals: Ragas · promptfooYou hold the keys

Public cloud models stay outside the perimeter. Connected on your decision and only through the masking gateway.

04
Where to draw the line

One platform, three places for the line

The choice depends on whose data you hold and whether you have a server room. The line can move later: the platform is the same.

Sovereign

Your building

Data never leaves the building. Your servers, inside your walls, every key in your hands. For banks, healthcare, government and anyone whose data must not travel.

Private cloud

Your country

Dedicated servers in a UAE data centre — Khazna, Equinix, Moro Hub — run by us. The same platform, no server room of your own. For companies without premises, small ones included.

Masked gateway

A shield before the cloud

When you need the strongest public model: a local gateway masks names, numbers and amounts before sending, restores them on the way back and refuses when unsure.

05
An honest comparison

Cloud API or your own platform

The cloud wins on time to start and access to the strongest models — and we do not hide it. For data, law, budget and control — the local platform. Sometimes the right answer is the cloud, and we will say so in the assessment.

CriterionCloud APIOn-premise platform
Where the data is With the provider, often abroad Inside your perimeter
Regulators: PDPL, CBUAE, DoH DPA, risk assessment, consents, approvals Localisation by construction
Cost as usage grows Grows with every request and seat Fixed: hardware + subscription
The very strongest models Available on release day Open ones trail by months
Time to start Minutes Weeks: hardware and setup
Tailoring to the company Prompts and RAG Fine-tuning on your data
Latency and offline The link and the provider’s queue Local network, works without internet
Dependence Prices, limits, model retirement — the provider decides A model swaps in a week
Audit trail of AI actions Provider-side, partial Complete, yours
Maintenance Not your concern In the subscription — but it is people and hours
06
Hardware

33 configurations — from a desktop box to a rack

Three typical answers are in the table. Below — a picker for your company: move the sliders, click the dots. You may buy the server yourself from any supplier; we specify, install and tune it. Prices are an October 2026 guide, VAT excluded; NVIDIA quotes change monthly.

SizeStaffModelsGPUServerPower
Starter
Budget alternative — 2× RTX 5090 (64 GB)
10–30Qwen3-32B · gpt-oss-120b · Falcon-H1-Arabic-34B1× NVIDIA RTX PRO 6000 Blackwell 96 GB — a workstation $20–23k0.6–0.9 kW, a normal socket
Business
2,000–3,000 tok/s on 70B, up to 7,000 on MoE; 50 agents at 32k context
50–200Llama 3.3 70B FP8 · Qwen3-235B INT4 — several in parallel4× RTX PRO 6000 Server (384 GB), a 4U server $105–120k3.4 kW, a rack and a 32 A line
Enterprise
~9,900 tok/s per server; 200–500 sessions
200+DeepSeek-V3.2 FP8 · Qwen3.5-397B · Mistral Large 38× H200 (1,128 GB) or 8× B200 (1,440 GB), a cluster $275–562k10–14 kW, colocation

Hardware map

Horizontal — server price, vertical — the largest model that fits. Dot size — how many people work at once at normal speed (≥15 tok/s). The hatched zone is your conditions. Hover or click.

People and speed

Left — how many people at once, right — one person’s speed on the recommended model. Configurations in ascending price.

Payback over 36 months

Cumulative cost of the selected configuration versus renting the same GPUs in the cloud and versus API fees at your token volume.

How we counted. One person’s speed is memory bandwidth divided by model size, scaled by 0.7 for NVIDIA and 0.85 for Apple; concurrency comes from vLLM/SGLang measurements (databasemart, 08.2026; SemiAnalysis InferenceX, 09.2026) and memory left for contexts. Spark, Mac mini and Mac Studio stacks follow Exo Labs and NVIDIA measurements: memory adds up, single-user speed barely grows because a token passes through every unit in turn. Mac mini M6 Pro and DGX Station GB300 are estimates from the 2026 line-ups. Rental is about 10% of the purchase price a month (RunPod, CoreWeave, 10.2026); electricity at the top DEWA tariff; maintenance 1.5% a year. API prices are blended (80% input, 20% output) as of October 2026. All of it is a guide, not an offer.

07
How we implement

Five steps: from assessment to handover

Every step ends with a tangible result and a decision — go further or not. The pilot is measured in minutes per task and in money, not in impressions. The first agent in production — within 10–14 weeks; timings are an estimate from published 2025–2026 rollout plans.

1–2 weeks

Discovery & use-cases

Interviews with leaders, inventory of systems and data, baseline metrics, network and security review. Together we pick the first three use-cases with a measurable effect.

→ Data-flow map, three use-cases, server spec and quote
2–3 weeks

Platform & first agents

We install the server in your building (yours or ours), deploy models, connect SSO, an isolated VLAN, connectors to mail, CRM and Jira, and load the documents.

→ The platform runs on your premises, three agents in staging
4–6 weeks

Pilot

Two weeks in shadow mode: the agent works alongside people, people grade it. Then 10–15% of live traffic and a pilot group of 20–50 staff. Metrics are written down in advance.

→ An accuracy scorecard, a go/no-go decision, an SLA
2–4 weeks per use-case

Production rollout

Full traffic with a rollback button, the personal CRM, training for all staff (5+ hours), champions in each department. Up to 5–10 agents within 3–6 months.

→ Agents in production, a runbook, a metrics report
subscription

Handover & care

All keys and documentation are yours. Monitoring, model updates, a new use-case every quarter, server access only with your approval.

→ A monthly hours-saved report, a growing catalogue of agents
95%
of GenAI pilots show no measurable effect — without a method
MIT NANDA, 2025
>40%
of agentic AI projects will be cancelled by 2027
Gartner, 2025
67%
success with a partner versus 33% going it alone
MIT NANDA, 2025
08
What it costs

Three packages. Prices published — competitors have none

No sovereign vendor publishes prices: Netveva, Cohere North, Mistral Enterprise, Aleph Alpha — all “on request”. We publish ranges. Two sums per package: a one-off rollout and a care subscription. Hardware is separate: flip the switch to see the total with it. Use is unlimited.

Currency

Starter

10–50 people: first agents and your documents at work
Implementation
$35–70k
Care per month
$3–5k
Hardware separately: $20–23k · timeline: 10–14 weeks
  • A 1× RTX PRO 6000 workstation (or 2× RTX 5090), models up to 120B
  • 3 agents: mail, documents, reports
  • Portal + Telegram / WhatsApp
  • Your documents indexed, answers cited
  • SSO and roles, action log
  • Team training — 1 day
most chosen

Business

50–300 people: agents per department, integrations and a personal CRM
Implementation
$90–200k
Care per month
$6–12k
Hardware separately: $105–120k · timeline: 3–5 months
  • A 4× RTX PRO 6000 server, models up to 70B FP8 and 235B MoE
  • 6–10 agents per role, task board, schedules
  • Microsoft 365 / Google, Jira, 1C, Bitrix24 — over MCP
  • Personal agentic CRM — included
  • No-code workflows with an approval step
  • PII masking, audit log
  • Training — 2 days, champions in departments

Enterprise

300+ people, regulated industries, several offices and countries
Implementation
from $200k
Care per month
from $12k
Hardware separately: $275–562k · timeline: 4–8 months
  • An 8× H200 / B200 cluster, high availability
  • Agents in every department, model fine-tuning
  • Arabic RTL, voice, mobile app
  • Compliance pack: PDPL, CBUAE, DoH, DIFC/ADGM — one-click report
  • A private cloud in the UAE instead of your own server room
  • 24/7 SLA, a dedicated engineer

Prices exclude VAT (5% in the UAE); dirhams at 3.6725. Hardware is quoted for purchase from a supplier, plus a 5% duty outside free zones. For private-cloud hosting instead of purchase — a monthly rental, calculated during the assessment. For banks and healthcare — a 15–30% compliance surcharge.

09
Personal CRM

A CRM the agents keep, not you

What it is

A CRM you do not have to fill in. The core is open-source (Twenty: PostgreSQL, built-in agents and MCP), the agents run on your local model. Contacts and deals appear from mail, calls and messengers by themselves; the database is yours, on your server, zero per-seat licences.

Migration
Contacts and deals from Excel, Bitrix24, amoCRM, HubSpot, Salesforce
Timeline
6–8 weeks for the base, 10–12 with telephony, 1C and forecasting
Price
$25–45k base standalone (extended — $45–90k); included in Business

What the agents do in the CRM

  • Contact and deal cards — from a mail, a call or a chat, no manual entry
  • Lead scoring by your criteria
  • Draft e-mails and WhatsApp replies in your tone
  • A deal brief before the meeting: history, risks, next step
  • Reminders and tasks that appear and close themselves
  • A sales forecast and a manager report on schedule
HubSpot Breeze · Salesforce AgentforceYour local CRM
Where data is processedAt OpenAI and the vendor’s partnersOn your server
Price$1 per lead, $0.50 per conversation — on top of the per-seat feeIncluded, no limits
Telephony and WhatsApp in the UAEVia third-party connectorse& / du, WhatsApp Business API — direct
Source code and databaseThe vendor’sYours

Salesforce State of Sales 2026: 54% продавцов уже работают с агентами, экономия на черновиках −36%

10
Consulting

Implement AI so it reaches the profit line instead of staying a pilot

88% of companies already use AI, yet only 39% see an effect on profit. The gap is not in models but in choosing tasks, data, rules and people. That is what we sell: assessment → strategy → rules → people → care. The platform joins at step two or three — as one option, not the goal.

ServiceDurationResultPrice
AI Pulse — rapid assessment2–3 weeksMaturity across six domains, risks, 10 use-cases with impact estimates and priorities, a hardware spec$8–15k
Use-case workshop1–2 daysLeaders and key staff; a ranked use-case list, an effect calculation, a pilot plan$5–10k
AI Blueprint — strategy & roadmap4–8 weeksStrategy, budget, KPIs, a 12-month quarterly plan, team requirements, a data model$25–60k
AI policy, PDPL & Dubai AI Seal3–6 weeksAn AI use policy, procedures, a risk register, UAE PDPL and AI Charter alignment, a Dubai AI Seal application$15–40k
AI Academy — training1 day · a 4–6-week programmeThree tracks: leaders ($6–12k a day), staff ($3–6k a day), IT. Hands-on with your own tasks, 5+ hours per person; a programme for a 50–250-person company — $15–40k$3–12k
Fractional Chief AI Officerper monthLeading the rollout: 2 days a month (advisory) or 1 day a week (embedded) — priorities, vendors, budget, board reporting$6–20k
The “Sovereign AI Launch” bundle4–5 monthsAssessment + strategy + policy + training + 3 months of fractional CAIO + a pilot of one agentic use-case. Hardware and subscription excluded$80–150k

Our day rate is $1,200–2,000: a strong boutique’s level without the Big Four mark-up, where an hour costs $400–800 — $3,200–6,400 a day. All package prices are AI Finder estimates from market corridors; the quote is fixed after the assessment.

295 000
companies Dubai plans to equip with agentic AI in two years
Dubai Media Office, 06.2026
57% → 7%
of UAE organisations adopted agents, but only 7% reached autonomous workflows
ServiceNow Index, 2026
$1 500
a year — the DIFC AI & Innovation licence with a 90% subsidy
DIFC, 2026
11
Questions

What CFOs and CIOs ask

What exactly leaves the company perimeter?

Nothing. Models, the document index, agents and the log run on your server. If a specific task needs a public model, data passes through a masking gateway: names, numbers and amounts are replaced before sending and restored after. This is enabled separately and only on your decision.

Will the model be outdated in six months?

Open models are updated every few months and the subscription covers it: we evaluate new releases on your tasks, size them to your hardware and apply them when you are ready. The hardware lasts for years: data centres depreciate GPUs over 5–6 years.

Who maintains the system?

We do, under the subscription: monitoring, updates, new use-cases, engineer support. Keys and access are yours; we connect only with your approval. One or two of your people are trained so day-to-day tasks do not need us.

Arabic and Russian?

Yes. For Arabic — Falcon-H1-Arabic (TII, Abu Dhabi) and Jais 2 (G42), with dialects and correct right-to-left layout. Russian and English are native to all major open models. Agents reply to customers in the language they wrote in.

How long until the first result?

The first agent in production — 10–14 weeks after the start; 9 weeks under ideal conditions. Starter in full — 10–14 weeks, Business 3–5 months to 5–10 agents, Enterprise 4–8 months including hardware lead time (6–14 weeks) and regulator approvals.

What if it does not work out?

That is why we start with a pilot whose metrics are written down in advance: minutes per task, accuracy, the share of mail handled without a person. The full-rollout decision is made on those numbers. The hardware stays yours either way and suits any other workload.

Do we need a server room?

Not for Starter: a workstation around a kilowatt and a normal socket. Business needs a cooled rack and a 32 A line; Enterprise needs colocation. With no room of your own, the platform goes into a private cloud in the UAE — data stays in-country.

How is this different from ChatGPT Enterprise or Microsoft Copilot?

Three things. Data does not go to the vendor. You pay for the rollout and a subscription, not $30–75 per seat — for 100 staff that is $36–90k a year. And the agents are built for your processes and systems, not for an average office: they work with your ERP, your contracts and your rules.

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.