Back to Catalog
Create Endpoint
thinkingmachines

Inkling-Small-NVFP4

Catalog model officially supported by Inference Endpoints.

This model is from our Model Catalog, and comes with pre-configured recipes. Deployment has been verified by Hugging Face.

/
$10 / h
per running replica
Nvidia H200
2x GPUs · 282 GB 46x vCPUs · 512 GB
$10 / h
Catalog Recipe
Pre-selected hardware for the current recipe.
  • Only you can access your endpoint, using a Hugging Face Token generated from your personal account.
Number of replicas
Automatically scale the number of replicas within Min and Max based on compute usage. Min is always 0 if Scale-To-Zero is active.
More options
Autoscaling Strategy
Control what type of trigger will cause your Endpoint to scale up.

This Catalog Recipe comes with a pre-configured vLLM engine.

This Catalog Recipe comes with pre-configured env values.

VPC Config
Check to activate and configure AWS PrivateLink