PacinfraX

AI CLOUD · GPU INFRASTRUCTURE

AI infrastructure,
on your terms.

Explore open-model services, GPU hosting and partner capacity. Choose how you build—with clearer ownership, deployment and cost boundaries.

OWN · PLACE · RUN
PacinfraXAI control plane

Models. Compute. Ownership.

Model APIs

Application layer

GPU fleet

Compute layer

Partner capacity

Infrastructure layer

Architecture concept, not a live network or capacity map.

BUILD YOUR WAY

Choose your deployment path.

Start with your workload and ownership model. Compare the requirements before making an infrastructure commitment.

DEVELOPER EXPERIENCE

Familiar code.
Clear boundaries.

Start with a familiar API contract. PacinfraX-specific controls stay explicit, versioned, and optional.

  • Streaming contract
    Inspect the event-stream request examples.
  • Request tracing contract
    Review request IDs without sharing prompt content.
  • Error contract
    Understand retry and configuration boundaries.
# Contract example; public endpoint is not enabled.
# Set PACINFRAX_API_KEY securely in your environment.
curl --no-buffer --fail-with-body --max-time 20 \
  --proto '=https' --tlsv1.2 \
  https://api.pacinfrax.ai/v1/chat/completions \
  -H "Authorization: Bearer $PACINFRAX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model-id>",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": true
  }'
# Prints raw server-sent events, not parsed assistant text.

Contract examples only: the public API endpoint is not enabled. Replace <model-id> with an authorized callable model ID before use. Examples print raw SSE events; they do not invoke a model on this page.

MODEL CATALOG

Find your model.
Understand the trade-offs.

Compare reference profiles and indicative token prices. Open a model passport for license, runtime and deployment requirements.

TEXT MODEL REFERENCE

GPT-OSS 120B

Input / 1M tokens
$0.13875
Output / 1M tokens
$0.555
Availability
Non-callable
View model passport
TEXT MODEL REFERENCE

Llama 3.3 70B Instruct Turbo

Input / 1M tokens
$0.962
Output / 1M tokens
$0.962
Availability
Non-callable
View model passport
TEXT MODEL REFERENCE

Qwen3.5 9B

Input / 1M tokens
$0.15725
Output / 1M tokens
$0.23125
Availability
Non-callable
View model passport

Indicative prices, not billable rates. Partial draft—not current capacity. License, runtime, region and commercial acceptance remain pending. Reference checked 2026-10-03. Explore all models and service units →

MODEL & SERVICE PRICING

Plan the whole cost.
Not just the tokens.

Compare model and developer-service estimates separately from GPU hosting. Understand the unit before choosing your deployment.

Indicative estimates only—not billable rates or a quote. Listed services are not available for invocation. Reference checked 2026-10-03.

Model inference
Separate input and output prices per million tokens. Compare selected model estimates above or review the full breakdown.
Developer services
Storage, sandbox CPU, sandbox RAM and interpreter sessions have separate units. CPU and RAM must both be included in sandbox estimates.
GPU hosting & infrastructure
Power, cooling, connectivity and managed operations need workload-specific terms; token prices do not include hosting.

TRUST BY DESIGN

Boundaries you can inspect.
Claims you can verify.

Understand project boundaries, credential handling and diagnostic practices before choosing your deployment.

Review security boundaries
  • 01
    Project-scoped API keys

    Expiry and revocation are part of the credential contract. Secrets are designed to be shown once.

  • 02
    Content-minimized telemetry

    Operational analytics use sanitized request metadata; prompt and output sharing stays opt-in.

  • 03
    Region-aware deployment planning

    Review required regions and data handling before committing to a deployment.

YOUR NEXT BUILD

Bring the workload.
We’ll make the path legible.

Prepare a workload brief locally before any capacity, region, price or timeline commitment. Customer intake is not enabled; nothing is submitted automatically.