scenario estimate

AI cost: model, RAG or local deployment

Compare three architectures under the same workload. A well-prepared knowledge base reduces context, enables smaller models and can make local deployment economically viable.

Large cloud model538,180 ₸per month

The full context is sent with every request

Smaller model + RAG12,453 ₸per month

Retrieval sends only relevant passages

Saving against the baseline: 98%
Local model450,000 ₸per month

Data and compute stay within your environment

Saving against the baseline: 16%
Calculation assumptions

30 days per month. Cloud without retrieval: $2 per million input tokens and $8 per million output tokens. Smaller model with RAG: $0.15 and $0.60 respectively. Indexing, storage and implementation are excluded.

Price does not determine quality. Before choosing a model, use a benchmark set to measure accuracy, answer completeness, citation quality and latency.

the right next step

First measure retrieval and model quality on your own documents. Then compare total cost of ownership, including infrastructure, support, security, knowledge-base updates and answer controls.