Pay for input, not output.
Billed per input token. Output is free: answers are read from the model's probabilities, not written as text. Figures may change before general availability.
What's included
The same service for every account, from the first decision. There is no plan to upgrade to.
- EU GPU regions
- GPU inference runs only in EU-RO-1, EU-SE-1 and EUR-IS-2.
- Packed and separate modes
- Every question shares one prefilled state in packed mode, the default. In separate mode each question is answered alone.
- Calibration published per model version
- Temperature scaling is fitted per primitive, and figures are published with the model version they belong to. No figures published yet.
- No state retained by default
- Counts, latency and token totals are kept so that usage can be billed.
- DPA on request
- The DPA names every sub-processor, including RunPod Secure Cloud. Request the DPA
- Email support
- Every account can write to the team directly. hello@hostjev.com
- Dedicated capacity
- A worker of your own or custom rate limits, arranged directly rather than sold as a plan. Talk to us