[Reference page](<https://kernodeck.com/en/usages/inference>)

# Choose the GPU for your inference application.

Batch embeddings, image ranking, or an interactive application: rent from Kernodeck the GPU suited to your own inference. Start with the 24 GB RTX 4090 at 110.00 USD for a lot of 1 GPU over 7 days. Choose the capacity based on your measured memory and your CUDA stack; you keep control of preparing and running the application.

- Batch or interactive processing
- Three budgets over 7 days
- Your engine and your settings

[Choose an RTX 4090 · 7 days](<https://kernodeck.com/en/compute/configure?config=rtx-4090&days=7>)

At Kernodeck, no ID and no KYC process on this GPU rental path. Account required: first name, last name, email and password; that does not mean anonymity.

Inference rental

One package for your application, a capacity chosen according to its workload: compare 24, 48, and 80 GB before renting.

## Key facts

| Fact | Detail |
| --- | --- |
| Period and unit | 3, 7 or 30-day plan for one lot. A lot contains 1 GPU, except B200: 2 GPUs, already included in its price. |
| Availability | Declared quantities per model in the offers. They do not reserve future availability or a delivery date. Hosting country and served region to be confirmed before purchase if your project requires them. |
| No KYC | At Kernodeck, no ID and no KYC process on this GPU rental path. Account required: first name, last name, email and password; that does not mean anonymity. |
| Setup | Ubuntu, PyTorch, Blender or a custom request. Versions, CPU, RAM, storage, network and access mode to be confirmed per project; their inclusion is not implied by the GPU model. |
| Pricing and terms | USD amounts for the lot and the entire period. Tax status and any fees to be confirmed in the applicable terms. |

## Available plans

Total price in USD for one lot and the full period. Memory is shown per GPU; declared stock is expressed in lots.

| Model | Memory per GPU | GPUs per lot | Duration | Lots | Total price | Declared stock in lots | Status | Configuration |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [NVIDIA GeForce RTX 4090 24GB](<https://kernodeck.com/en/compute/rtx-4090>) | 24 Go | 1 | 7 days | 1 | 110.00 USD | 265 | For batch processing whose workload fits within 24 GB per card. | [Configure this plan](<https://kernodeck.com/en/compute/configure?config=rtx-4090&days=7>) |
| [NVIDIA L40S 48GB](<https://kernodeck.com/en/compute/l40s>) | 48 Go | 1 | 7 days | 1 | 148.00 USD | 74 | For examining an interactive or multimodal workload with 48 GB per card. | [Configure this plan](<https://kernodeck.com/en/compute/configure?config=l40s&days=7>) |
| [NVIDIA H100 SXM 80GB](<https://kernodeck.com/en/compute/h100-sxm>) | 80 Go | 1 | 7 days | 1 | 475.00 USD | 63 | For a memory requirement beyond 48 GB, with a qualified CUDA stack. | [Configure this plan](<https://kernodeck.com/en/compute/configure?config=h100-sxm&days=7>) |
| [NVIDIA B200 SXM](<https://kernodeck.com/en/compute/b200>) | 180 Go | 2 | 7 days | 1 | 2,071.00 USD | 4 | Two 180 GB GPUs each for an application that organizes their use. | [Configure this plan](<https://kernodeck.com/en/compute/configure?config=b200&days=7>) |

## For your batch processing: start from the need, then the package

Do you want to produce embeddings for a collection of documents or classify a set of images? The RTX 4090 is a candidate at 110.00 USD for a lot of 1 GPU with 24 GB over 7 days if your CUDA pipeline fits within that memory along with its inputs, its temporaries, and a measured margin. This package frames your hardware budget; you set the number of files, the recovery rule, and the results to keep.

If your workload's memory peak exceeds this capacity, compare the L40S: 148.00 USD for a lot of 1 GPU with 48 GB for 7 days. It provides more room per card for the load you have measured. The choice between the two comes down to that capacity and the compatibility of your software stack, not to a promised document volume.

For independent partitions, the B200 offers a different organization: 2,071.00 USD for a lot of 2 GPUs with 180 GB each for 7 days, with both cards already included in that price. Compare this option when your application can distribute processing across multiple workers. It does not merge the two memories into a single 360 GB space.

- [Choose 7 days on RTX 4090 — 110.00 USD](<https://kernodeck.com/en/compute/configure?config=rtx-4090&days=7>)
- [Choose 7 days on L40S — 148.00 USD](<https://kernodeck.com/en/compute/configure?config=l40s&days=7>)
- [Review the B200 lot for distributed processing](<https://kernodeck.com/en/compute/b200>)

## For an interactive application: choose the memory for your full workload

An assistant or an internal tool must account for concurrent requests, their length and engine-specific caches. The L40S at 148.00 USD for a lot of 1 GPU with 48 GB for 7 days is a candidate if this full workload fits in its memory and if your engine supports its CUDA stack. Your measurement of the queue and of the responses remains the criterion for selecting the configuration.

If you need more memory on a single card, compare the H100 SXM: 475.00 USD for a lot of 1 GPU with 80 GB for 7 days. It adds capacity per GPU; it is not enough to guarantee your application's latency or throughput. If your requirement already fits in 24 GB, the RTX 4090 remains worth comparing before committing to a higher-tier plan.

This offering is a GPU rental for your application. A managed inference API is a different category of service: it may suit you if you are first and foremost looking for an endpoint operated on your behalf. With Kernodeck, you provide your engine, its dependencies and its operation; the preparation and access arrangements required for your project remain to be confirmed before you choose your environment.

- [Choose 7 days on H100 SXM — 475.00 USD](<https://kernodeck.com/en/compute/configure?config=h100-sxm&days=7>)
- [Compare the L40S and its plans](<https://kernodeck.com/en/compute/l40s>)
- [Go back to 24 GB with the RTX 4090](<https://kernodeck.com/en/compute/rtx-4090>)
- [Prepare your application's software stack](<https://kernodeck.com/en/docs/environnements-reproductibles>)
- [Compare the 3, 7 and 30-day plans](<https://kernodeck.com/en/pricing>)

## Your plan, payable directly in crypto.

Pay for your Kernodeck rental directly in crypto, with no mandatory top-up: BTC on Bitcoin and USDT on Tron (TRC-20), among others. The process requires neither an ID document nor a KYC procedure; an account with first name, last name, email and password is required, with no promise of anonymity.

- [Compare the eight accepted assets and networks](<https://kernodeck.com/en/docs/crypto-payments>)

## Define the deliverable that justifies your rental

For file processing, define the volume to process, the output format, and the resume rule. A finished file must be recognizable without re-reading the entire run. For an interactive service, define the maximum request size, the acceptable latency, and the behavior when capacity is reached. The same model can require two very different organizations depending on these constraints.

Keep a representative sample with short, typical cases and cases close to your limits. In an embeddings pipeline, associate each vector with the identifier and version of its input. In text generation, record the generation settings used for evaluation. You must be able to explain a difference in results without immediately attributing it to the GPU.

## Choose the memory for the entire inference workload

The size of the model's files does not describe all the memory used during inference. You must also account for temporary tensors, inputs, retained outputs, and, for the models concerned, the attention key-value cache. This cache can become significant when sequences grow longer or when multiple requests are processed together.

Start with a card whose memory allows your representative trial with a measured margin. Models of 24, 32, 48, 80 GB and beyond meet different needs; no capacity guarantees that a given model will fit with all its settings. Reducing precision or quantizing can change the footprint, but requires verifying software support and output quality on your own inputs.

- [Compare the catalog's memory capacities](<https://kernodeck.com/en/compute>)
- [Identify the phase consuming memory](<https://kernodeck.com/en/docs/observer-un-lancement>)

## Pick the GPU compatible with your software stack

List the inference engine, its specific operators, and the extensions you depend on before choosing the hardware. A PyTorch application can offer several execution paths, whereas a specialized extension supports only one. Verify the full chain on CUDA for NVIDIA or on ROCm for AMD, including model loading and its preprocessing.

Keep a minimal command that runs the pipeline all the way through to writing a result. Only then enable your optimizations one at a time. Each change of precision, compilation, or engine must preserve a comparable quality check. A successful load demonstrates that the weights are readable; it does not demonstrate that all the necessary computation paths work.

## Set the batch according to your processing goal

Example method: build three groups of texts by length, then process each with a batch of 1, 2, and 4 inputs. These sizes serve as trial points, not a universal recommendation. For each combination, note the completed inputs, the total time, the maximum observed memory, and the errors. Stop the progression when a limit appears instead of hiding failures in an average.

For an interactive service, add the time spent in the queue. A larger batch can change the throughput and the latency felt by a request; a single average is not enough to choose. For offline processing, make sure that grouping does not mix up the order of the results. Ultimately choose a configuration that respects your quality criterion and your latency constraint.

## Move to multiple GPUs if your application can distribute the work

If each instance of the model fits on one card, you can organize several workers that consume distinct partitions of the inputs. You must then coordinate identifiers, resumes, and output collection. If the model must be split across cards, use a parallelism strategy supported by your engine and verify its communication requirements.

Ordered lots describe the quantity of hardware, not the application batch or a merged memory space. For B200, one lot includes two cards; for the other offerings, one lot includes one card. Document in your file the number of workers planned, the share of inputs assigned to each, and how you will verify that a job is actually finished.

## Choose a period that includes the checks and the export

For a first 3-day rental, set a limited goal: install, validate the pipeline, and produce a first usable result. A 7-day duration can be used to explore more variants; 30 days to repeat a process and consolidate its operation. These are ways to organize the work, not promises of execution time.

On exit, keep the weights or their version, the configuration, the quality checks, the measurements actually obtained, and the exported results. You choose your software and processes autonomously; Kernodeck does not inspect their content. Prepare your own access, backups, and the authorizations required to use the models and data.

## Questions before purchase

### Which plan should I choose for a first inference workload?

If your CUDA pipeline and a representative batch fit in 24 GB with a measured margin, start the comparison with the RTX 4090: 110.00 USD for a lot of 1 GPU for 7 days. The plan covers that rental period, without fixing the number of files your application will be able to process. Keep installation, output validation and export in your schedule; the 3 and 30-day plans let you choose a different period.

### When should I pay more for an L40S or an H100 SXM?

Compare the 48 GB L40S when the full workload exceeds the margin available on 24 GB, then the 80 GB H100 SXM if you want more memory on a single card. The 7-day plans cost 148.00 USD and 475.00 USD respectively, each for a lot of 1 GPU. Include the model, the inputs, the temporary files and the caches in your measurement; higher capacity alone does not guarantee a better response for your application.

### Does the rental include a ready-to-call inference API?

The offering presented here is a GPU rental for your own application, and does not include a managed inference API advertised as ready to call. You choose the engine, the models, the dependencies and the settings, then you organize their operation. If your priority is a managed service, compare that category of offering separately. To rent with Kernodeck, have the preparation and access arrangements required by your project confirmed.

### Is the B200 the logical choice for multiple simultaneous workloads?

It becomes an option worth considering if your application can assign independent partitions to multiple workers or supports explicit distribution. At 2,071.00 USD for 7 days, a B200 lot already includes 2 GPUs with 180 GB each. The memories do not merge automatically; a distributed model requires a strategy compatible with your engine. Also compare the cost of a single card if it is enough for the planned workload.

## Choose and configure

- [Prepare a reproducible inference environment](<https://kernodeck.com/en/docs/environnements-reproductibles>)
- [Consider a 48 GB NVIDIA option](<https://kernodeck.com/en/compute/l40s>)
- [Consider an AMD option for a ROCm pipeline](<https://kernodeck.com/en/compute/mi300x>)
- [Choose a duration of 3, 7, or 30 days](<https://kernodeck.com/en/pricing>)
- [The eight accepted crypto pairs](<https://kernodeck.com/en/docs/crypto-payments>)
- [Choose a capacity for inference](<https://kernodeck.com/en/usages/inference>)
