EU-hosted AI inference for high-frequency tasks

Sval Compute is an inference API serving models for tasks such as classification, moderation, embedding, and reranking. Zero data retention. Full infrastructure in the EU.

Models for task calls

Smaller models are well suited for offloading specialized tasks in agentic systems: they respond faster, cost less per call, and hold up at high call volumes. Sval serves this tier.

EU-only inference

Model inference runs on GPU hardware we own and operate in Sweden. The entire serving path, from network entry to model output, stays in Sweden under EU jurisdiction.

Zero data retention

Prompts and completions exist in memory only, while being processed. The prefix cache is memory-only with short TTL and per-tenant isolation.

Products

Small models for high-frequency tasks

Classify

LIVE

A classification endpoint for content policy checks. Send it text, get back a label from a fixed set of safety categories. Typical usecases are screening user-generated content, checking prompts and model output in an LLM application, and flagging items for human review. The calls are short and cheap, which matters when volume is large.

Currently serving granite-guardian-3.3-8b (Apache-2.0). Check our Live catalog.

Retrieve

LIVE

Embedding and reranking endpoints, for search over your own documents. Embedding turns text into vectors so you can search by meaning rather than by keyword. Reranking scores the candidates that search returns and orders them by how well they match the query. Together they are the retrieval half of a RAG pipeline, and both run at volume.

Currently serving qwen3-embedding-4b and qwen3-reranker-0.6b (Apache-2.0). Check our Live catalog.

Generate

EVALUATING

Extract fields, transform, summarize, translate, draft, label against your own categories. These calls run thousands of times a day, so cost per call is what matters, and a small model does the work well. Optionally, a JSON schema can be sent in the request to constrain the structure of the model output.

We are currently evaluating candidates like gpt-oss-20b (21.5B, Apache-2.0) and mistral-nemo-instruct-2407 (12B, Apache-2.0).

Contact

Get access, or just say hello

Sval Compute is a Swedish inference provider, built and operated by Albin Pansell and Simon Riis. API access is available on request: write to us with a few lines about what you want to run, and we will get back to you with access details. If there is a model you would like us to host, name it and we will look into it. Questions and business inquiries reach the same address.

hello@svalcompute.com