Application Performance Tuning

Job Count

To achieve the best throughput of cryptographic jobs (such as Sign or Decrypt) in your application, arrange for multiple jobs to be on the go at the same time, rather than doing them one at a time. This is true even when using only a single HSM in your system.

When using an nShield HSM, Entrust recommend that you set the number of outstanding jobs within the rec. queue (recommended queue) range specified by the enquiry output for the module.

If you are sending single jobs synchronously from each thread of your client application, try to keep the number of threads within this rec. queue range for best throughput.

When using higher-level APIs, such as PKCS#11, your application could benefit from increasing the thread count above the rec. queue range or the number that gives the best throughput when using nCore directly.

If you are load-balancing across multiple HSMs and want to maximize throughput across all of them, then use the sum of all rec. queue ranges for each of the modules to set the target for the outstanding jobs.

The ncperftest utility supports performance measurements of a range of cryptographic operations with different job counts and client thread counts. You may find this useful to inform tuning of your application. Run ncperftest --help to see the available options.

Client Configuration

If your application is coded directly against nCore, you have a choice of sending multiple jobs asynchronously from a single client connection to the hardserver, or having multiple threads each with their own client connection to the hardserver with a single job sent synchronously in each. You can use the --threads parameter to the ncperftest utility to experiment with the performance impact of having more threads/connections with fewer jobs outstanding in each, or having fewer or just one thread/connection with more jobs outstanding in that connection.

When using higher-level APIs such as PKCS#11, all cryptographic operations are synchronous, so larger numbers of threads must be used to increase the job count and make full use of HSM resources. These APIs automatically create a hardserver connection for each thread. If many HSMs are being used, a great many threads may be required to achieve best throughput. You can adjust the thread counts in the performance test tools for these APIs (for example, cksigtest for PKCS#11) to gauge how much concurrency is required for best throughput in your application.

Highly Multi-threaded Client Applications

If your application is highly multi-threaded, operating system defaults may not be optimal for best performance:

You may benefit from using a scalable memory allocator that is designed to be efficient in multi-threaded applications, examples include tcmalloc.

On some systems the default operating system scheduling algorithm is also not optimized for highly multi-threaded applications. A real-time scheduling algorithm such as the POSIX round-robin scheduler may yield noticeable performance improvements for your application.

File Descriptor Limits (Linux)

On Linux systems, large numbers of threads each with their own hardserver connection will require your application to make use of large numbers of file descriptors. It may be necessary to increase the file descriptor limit for your application. This can be done using ulimit -n NewLimit on most systems, but you may need to increase system-wide hard limits first.

Cryptographic Acceleration Profiles

nShield 5 HSMs support configurable FPGA acceleration profiles that control the balance of hardware acceleration between traditional public-key (PK) and post-quantum (PQ) algorithms.

If you use post-quantum algorithms (ML-DSA/ML-KEM), you should select the HY (Hybrid) profile. The HY profile adds FPGA acceleration for ML-DSA (from firmware v13.8.4) and ML-KEM (from firmware v13.8.6). This improves ML-DSA signing throughput by approximately 7x and verification by approximately 3x compared to the default PK profile.

The tradeoff is reduced throughput for some traditional algorithms. RSA signing throughput is approximately halved, and throughput for ECDSA and DH operations over larger curves (P-384 and P-521) is reduced by approximately 40% to 50%. Operations over smaller curves (P-256) are less affected. Both profiles support the same set of algorithms — only the performance balance differs. You can use the perfcheck and ncperftest utilities to measure the performance of specific algorithms on your HSM with either profile.

Profile ID Profile Name Description

PK

PK Algorithm Acceleration

Maximises acceleration of traditional public-key algorithms (RSA, ECDSA, ECDH, DH, DSA). This is the default profile.

HY

Hybrid PK and PQ

Adds FPGA acceleration for post-quantum algorithms (ML-DSA, ML-KEM) while retaining acceleration of traditional algorithms at reduced throughput.

  • Only Mid and High-speed nShield 5 variants support acceleration profiles. Base-speed variants do not have this option.

  • You can only change the profile when the HSM is in maintenance mode. A module reset is required for the new profile to take effect.

  • Changing the profile does not affect the Security World configuration.

  • Different HSMs in the same Security World can use different profiles independently.

  • The default profile is PK.

nShield 5s (PCIe)

Use the hsmadmin select acceleration command to manage acceleration profiles on the nShield 5s.

To view the available profiles and the currently loaded accelerator:

hsmadmin select acceleration --show

To select the HY (Hybrid) profile:

hsmadmin select acceleration --set HY

To revert to the default PK profile:

hsmadmin select acceleration --set PK

Place the module in maintenance mode before changing the profile. Use nopclearfail --maintenance to enter maintenance mode, then reset the module with nopclearfail --clear or hsmadmin reset for the new profile to take effect.

nShield 5c (network-attached)

Use appliance-cli from a privileged client to manage acceleration profiles on the nShield 5c.

To check the currently loaded accelerator:

appliance-cli -m 1 gethsminfo | grep loaded_accelerator

Example output:

            "loaded_accelerator": "HY"

To select the HY (Hybrid) profile:

appliance-cli -m 1 sethsmoption --option acceleration --value HY

To revert to the default PK profile:

appliance-cli -m 1 sethsmoption --option acceleration --value PK

Use -m MODULE_NUMBER to target a specific module if more than one is installed.

Place the module in maintenance mode before changing the profile. Enter maintenance mode using maintmode via the serial console or front panel. After changing the profile, reset the module with nethsmadmin -m 1 -r for the new profile to take effect.

On nShield 5c image versions before v13.9.7, appliance-cli gethsmoption --option acceleration disrupts the module by temporarily switching it to maintenance mode and back. Use gethsminfo as shown above to query the current profile without disruption.