Intelligence Suite
EN

Platform & Technology

Everything between your code and the bare metal.

One connection gives your enterprise access to every model, while Indosat manages the complexity underneath: gateway, cache, router, meter and GPUs.

The architecture layers

The full AI stack, from power to applications

The combination decides the cost. Indosat runs the stack for local models and connects you to global ones.

  1. 06ApplicationYour applications and agents, plus ready-made industry solutions for learning, customer care and fraud intelligence.
  2. 05Optimized operations & governanceOrchestration, access control, rate limits, spend controls, token metering and usage reporting.
  3. 04ModelsOpen-source and closed-source models, Western and Eastern models and Sahabat-AI LLM.
  4. 03Inference platformModel serving at scale, Model-as-a-Service, and multi-model, multi-modal delivery through one API.
  5. 02Compute (GPU) infrastructureGPUs and TPUs with high-speed networking and storage.
  6. 01Data centres, power & networkHybrid, modular data centres, with land, water, power and Indosat's national network.

Optimization

Repeat requests, returned in milliseconds.

Traditional request

Compute every time.

Every request travels through the full model computation path, increasing latency and cost.

Cached request

Reuse what’s ready.

An exact cache hit returns the saved response directly from the gateway. No model call, no model API cost.

< 10 ms

to return an exact cache hit from the gateway, with no model call

Performance

Maximum tokens per second from every GPU.

Smarter model partitioning
Large models are split into expert parts spread across GPUs on NVLink, so each GPU holds only what it needs. Prefill and decode run on separate GPU pools for faster first tokens.
Quantization (AWQ / FP8)
FP16 → FP8 or INT4 cuts the vRAM footprint 2–4× with minimal impact on reasoning accuracy.
Continuous batching, PagedAttention, speculative decoding
Up to 300% higher throughput, 4× larger batches, and around 2× output speed.

Indosat Locally Hosted Models

The right model for every need. Scalable. Secure. Sovereign.

  • Nano / Edge

    1B

    Lightweight & on-device

    Fast, low-latency AI for real-time and simple tasks.

  • Small / Efficient

    7B+

    Efficient & cost-effective

    Everyday enterprise tasks and automation.

  • Enterprise

    30B

    Enterprise-grade intelligence

    Complex workflows and domain use cases.

  • Near frontier

    70B

    High-performance AI

    Advanced reasoning for specialist use cases.

  • Frontier

    100B+

    Best-in-class intelligence

    Research, innovation and critical scenarios.

See the platform on your own workload.

One API in. Governed intelligence out. Nothing to manage in between.