>

AI consulting services: deciding what to build, what it costs to run, and how to keep it safe

Most AI consulting services will help you build something. Fewer will say which of five quite different approaches fits your problem, what it costs in year two when adoption has tripled, or what has to be true about your data first. This is the decision layer of our enterprise AI implementation practice. If you know what you want built, our agentic AI, enterprise GPT and data science pages go straight to the mechanics. If a pilot demonstrated well and then stalled, start here.

5
Approaches we compare
Eval-first
Threshold set pre-go-live
Year 2
The budget that matters
CMMI 3
Appraised delivery

What we settle before anything is built

  • Which of prompt, RAG, fine-tune, agent or classical ML fits
  • Whether your master data can support the answer at all
  • The test set and the accuracy threshold for go-live
  • Token, vector-storage and inference cost at year-two volume
  • Scoped credentials and an approval gate on irreversible actions
  • What data may go to which model, and how long it is kept
The uncomfortable first finding

Enterprise AI projects fail on data, not on models

The model is rarely the constraint any more. Foundation models on Azure OpenAI, AWS Bedrock and Google Vertex are good enough for most enterprise tasks, and the gap between the best and the third-best is smaller than the gap between clean and dirty master data. What stops projects is duller: the same customer exists three times under different identifiers, the material master holds free text where it should hold attributes, two systems disagree about which order is open, and nobody owns the discrepancy.

A model reading that inherits the mess and presents it fluently. Inconsistent data in a spreadsheet looks inconsistent; the same data in a generated paragraph reads as authoritative. So the honest first step is often data work — golden-record definitions, deduplication rules, an agreed source of truth per entity. If that is not fundable, the right recommendation is a narrower use case that does not depend on it.

Two other patterns recur. Pilots never designed to be evaluated, where the acceptance criterion quietly became "it seems good". And cost surprise: AI running cost is usage-based, so it is lowest at exactly the moment you decide whether to proceed. Forty pilot users asking short questions tell you little about the bill when eight hundred people use it daily with long retrieved context and multi-step agent loops.

The pilot that demonstrated wellNo evaluation set, so nobody can show whether it improved or degraded after the first model change
Confidently stale answersRetrieval over a knowledge base nobody has maintained for two years, returned as current policy
One service account for every agentAn agent authenticated as an integration user inherits every permission that user ever needed
Year-two invoice shockBudget approved on pilot token volumes, then adoption triples and the vector index doubles
What the practice does

The decisions, and the engineering that follows

We work on platforms you have already licensed: Azure OpenAI and Copilot Studio, AWS Bedrock, Google Vertex, Snowflake and Databricks for the data layer, Salesforce Agentforce and Einstein, ServiceNow Now Assist, SAP BTP AI services, and Workato Agent Studio with Agent Guardrails. Delivered under our CMMI Level 3 appraised framework with ISO 27001 certified information security.

Generative AI consulting and use-case triage

The first job is to throw things out. Most candidates fail one of three tests: no measurable baseline, the answer exists in no system you control, or a regulator will want a named human accountable.

  • Baseline captured first — handling time, error rate, volume, cost per case
  • Feasibility screen against where the source data actually lives
  • A written rejection list, so nothing is re-proposed each quarter
  • Sequencing so use cases sharing a data foundation go together

RAG implementation — and when retrieval is wrong

Retrieval suits questions answerable from documents you control and keep current. It is wrong when the corpus is stale: RAG does not fix stale content, it makes it sound authoritative and attaches a citation.

  • Corpus audit first: owner, last-review date and retirement path per set
  • Chunking, hybrid search and reranking tuned against your own questions
  • Permissions enforced at retrieval by user identity, not by prompt instruction
  • Explicit refusal when the corpus does not hold the answer

Agentic AI development with scoped credentials

An agent that can act has authority, and authority needs a boundary. On a shared service account the blast radius is unbounded: one injected instruction reaches everything that account can touch.

  • One scoped token per agent, granting only the actions it needs
  • Read freely; write only through whitelisted idempotent, validated tools
  • A replayable audit log of every action and its inputs
  • Human approval on the irreversible — payment, commitment, deletion

LLM integration services and AI in ERP, CRM and ITSM

Integration is where generative AI stops being a demonstration. The capability belongs inside the workflow people are already in — invoice extraction landing in the ERP behind a confidence threshold, summarisation inside the ticket or requisition itself — not a chat window they must remember exists.

  • Azure OpenAI and Copilot Studio where identity already lives in Entra
  • AWS Bedrock or Google Vertex where the data sits there and egress matters
  • Salesforce Agentforce, ServiceNow Now Assist and SAP BTP AI in-platform
  • Workato Agent Studio with Agent Guardrails when a workflow spans systems

The data layer: Snowflake, Databricks, master data

Before promising a forecast or an extraction accuracy we check whether entity definitions are stable, history is long enough to hold seasonality, and the label you want is recorded rather than reconstructed.

  • Golden-record and survivorship rules for customer, material and vendor
  • Feature and embedding pipelines versioned alongside the model
  • Lineage from every answer back to the source record
  • A blunt verdict, and a narrower alternative, when data cannot support it

Evaluation and MLOps services after go-live

A model in production has a decay curve. Prompts change, vendors deprecate model versions, documents change, and real inputs drift from what you tested. MLOps here is the machinery that notices.

  • A defined test set scored on groundedness, accuracy and correct refusal
  • Regression run on every prompt, model-version or tool change
  • Drift monitoring on inputs, output mix and accuracy once ground truth lands
  • Named ownership and an agreed re-indexing or retraining trigger

AI governance, the EU AI Act and India's DPDP Act

Governance fails when it is a policy nobody can apply to a specific decision. What people need is a list: these tools are approved, this data class may go to that model, prompts are kept this long.

  • An approved-tools register that names the shadow AI in use today
  • Data classification mapped to models — what may leave your tenancy
  • Retention and logging rules for prompts and outputs holding personal data
  • EU AI Act risk classification; DPDP notice, purpose limitation, consent
Choosing the approach

Which AI approach fits the problem

ApproachBest forData neededRunning cost profileMain risk
Prompt / assistantDrafting, rewriting, summarising text a person already has openNone beyond the request itselfLowest per use; scales with adoption and conversation lengthNo grounding — plausible rather than sourced, and over-trusted
RAG (retrieval)Questions answerable from documents you control: policy, product, contractsA current corpus with owners and access-control metadataTokens driven by retrieved context size, plus storage and re-indexingA stale corpus answered confidently, with a citation attached
Fine-tuningConsistent tone, strict output format, one narrow repetitive taskHundreds to thousands of curated input/output pairsTraining cost each time, then cheaper per call than long promptsTeaches style, not facts — teams expect knowledge and do not get it
Agentic workflowMulti-step work across systems where the path varies per caseDocumented rules with exceptions, scoped API access, historical casesHighest and least predictable — many calls per case, plus retriesAuthority: wrong actions at machine speed on too broad a credential
Classical MLForecasting, propensity, anomaly detection, pricingYears of consistent labelled history with stable definitionsCheapest and most predictable at inference; cost sits in the pipelineSilent drift, and use where a known stable rule would do

Most production systems combine two or three of these. The discipline worth keeping is naming which approach owns which part of the problem, because each fails differently and each is monitored differently.

How we deliver

How an enterprise AI engagement runs

The order matters more than the labels. Evaluation is designed before the build, and the cost model uses second-year volume rather than pilot volume.

01

Triage and baseline

We cut the candidate list. Each survivor gets a recorded baseline; each rejection a written reason. Without a baseline, "it feels faster" becomes the only evidence available later.

02

Data readiness verdict

Whether the entities, history and labels the use case depends on exist and agree. This is the stage that most often changes the plan.

03

Evaluation set, before the build

Real cases with agreed correct outcomes, a scoring method, and a go-live threshold signed off in advance. Fix it afterwards and it drifts to wherever the model happened to land.

04

Thin slice in your own tenancy

One narrow end-to-end path, wired into the system of record, with scoped credentials and the audit log in place from the first commit rather than before the security review.

05

Shadow run and the threshold decision

Live cases in parallel, humans still deciding, results scored against the test set and the baseline. Missing the threshold and not going live is a legitimate outcome.

06

Governance handover and cost review

Named owner, approved-tools entry, retention rules, regression suite, drift alerts — then a review once adoption has grown, because the bill at real volume is the number the business case guessed at.

Related

Where to go for the mechanics

Questions we get

AI consulting, answered without the hedging

Why do enterprise AI projects fail on data rather than models?
Because foundation models are already good enough for most enterprise tasks, while master data frequently is not. If the same customer exists three times, or the material master holds free text where it should hold attributes, the model inherits that and presents it fluently. Bad data in a spreadsheet looks bad; in a generated paragraph it looks authoritative. So the honest first step is usually data work, not model work.
When is RAG the right approach, and when is it not?
RAG suits questions answerable from documents you control and keep current — policies, procedures, product and contract content. It is wrong when the corpus is out of date, because retrieval does not fix stale content. It makes a stale answer sound authoritative and attaches a citation, which is worse than no answer. Audit the corpus and assign owners before building the pipeline.
Can an AI agent use our existing integration service account?
It should not. An agent on a shared account inherits every permission that account has accumulated, so one reasoning error or injected instruction has an unbounded blast radius. Each agent gets its own scoped token for only the actions it needs, every action is written to a replayable audit log, and anything irreversible passes a human approval gate.
How do you decide whether an AI system is good enough to deploy?
Against a test set built before the build: real cases with agreed correct outcomes, a scoring method, and a threshold signed off by the process owner in advance. Fixing the threshold before anyone sees the model perform is deliberate, because afterwards it drifts to wherever the model landed.
What drives the running cost of enterprise AI?
Three things: tokens per interaction, which retrieved context length and conversation depth dominate; vector storage plus re-indexing as content changes; and inference capacity if you host open-weight models. All of it is usage-based, so it is lowest during the pilot. Budget against second-year volume with realistic adoption, not against the pilot invoice.
Where does AI genuinely pay off in an enterprise today?
Document extraction and classification into ERP, demand forecasting measured against a recorded planner baseline, code and test generation, service-desk deflection on questions the knowledge base really answers, and summarisation inside a workflow people already use. Fully autonomous decision-making in finance or compliance is not there yet for most organisations, and we will say so rather than sell it.
What does AI governance need to contain to be usable?
Four concrete things rather than a policy statement: an approved-tools register, a data classification saying which class of data may go to which model, retention and logging rules for prompts and outputs, and named ownership per system. From those you can answer a specific question at a specific moment, which is what governance is for.
How do the EU AI Act and the DPDP Act affect what we build?
The EU AI Act classifies systems by risk, so obligations follow the use case — employment, credit and biometric uses carry materially more than a drafting assistant. India's DPDP Act bites on personal data in prompts, outputs and logs: purpose limitation, notice and retention apply to a prompt log as much as to a database.

Let's build what's next — together.

Whether it's setting up your India GCC, modernizing your enterprise stack, or hiring 50 engineers in 30 days — we'd love to scope it with you.