Successfully Testing KernelCare Live Patching: Best Practices for Administrators

A robust test for KernelCare Live Patching verifies more than just a successful patch download: The running kernel must be supported, the patch status must be verifiably active, and the application must remain stable under a realistic load cycle. Start on a production-like staging host, then roll out through QA and Canary, and document termination criteria. Live patches delay reboots, but they do not replace them. Therefore, continue to schedule regular kernel updates and reboots as an integral part of your operations.

Understanding KernelCare Livepatch Correctly

KernelCare is TuxCare's agent for Kernel Live Patching on supported Linux systems. It applies security fixes to the running kernel without requiring an immediate server restart. Whether a patch is applicable depends on the specific combination of kernel build, distribution, and architecture; the availability of an agent package alone does not guarantee this support.

From a technical standpoint, the Upstream Linux Livepatch Framework describes a consistency transition in which affected tasks safely switch to modified code. This documentation explains the general kernel framework, but not necessarily the implementation approach of every KernelCare variant. For product-specific features and operational decisions, therefore, the information provided by TuxCare is decisive.

A patch that has been downloaded or reported as applied initially verifies the functionality of the patch chain. It does not prove that database connections, storage accesses, network paths, batch jobs, and business transactions will remain error-free under actual load. A robust test therefore evaluates patch status, system metrics, and application results collectively.

Regular Kernel Updates are still required. Live patches do not modify the installed kernel package and do not automatically cover hardware support, functional changes, or all driver adjustments in a new kernel. Furthermore, TuxCare provides patches for a specific kernel only as long as its vendor releases security updates for that series.

KernelCare also affects the kernel and must be distinguished from userspace patching. A successful test does not confirm either a LibCare patch status or the complete resolution of all vulnerabilities on the host. Live patching thus complements package management and change management: It can apply urgent kernel fixes more quickly, while regular package updates and scheduled reboots remain part of the maintenance strategy.

Components, Platforms, and Clear Distinctions

Before the test, the TuxCare architecture must be completely separated. The KernelCare agent runs on the target host, retrieves patch sets, and applies them to the running kernel. ePortal In contrast, it is an optional, self-managed component for the centralized control of patch sources and rollouts, for example in controlled or isolated networks. Both components perform different tasks and are not interchangeable.

Separate from this, LibCare is available as an add-on for userspace components such as glibc or OpenSSL. A successful KernelCare test does not check either the installation or the patch status of LibCare. Test logs should therefore record these levels separately: the kernel patch status, central distribution, and userspace patching each require their own documentation, approvals, and, if necessary, their own staging systems.

The first practical task is to create a robust inventory. This should include the distribution and release, the kernel that is actually booted, the architecture, the type of virtualization, enabled security mechanisms, and installed kernel modules. Equally important are storage and network drivers, as well as security, backup, and monitoring agents. These characteristics determine whether a staging host realistically represents the future production environment and whether the patch being applied is compatible with the kernel build.

The final decision regarding support is not made by a general distribution list alone. Check the specific combination of distribution, kernel version, and architecture in the TuxCare compatibility and patch database. Only this check distinguishes an installable agent from a kernel that is actually supported. It should be documented before any rollout planning and repeated whenever the kernel is changed.

Secure Boot constitutes its own platform class. The agent requires a suitable chain of trust for its kernel modules. TuxCare specifies Agent version 3.0-2 as the minimum version required for the automated Secure Boot process on supported RPM systems; this specification is not a general minimum version requirement for KernelCare and does not apply to manual MOK registration. The automated process requires, among other things, EFI boot, shim, and enabled Secure Boot, and is not intended for Debian or Ubuntu. Therefore, a scheduled reboot is part of the validation process for this configuration.

Before installation, you should also check for any existing live-patching services. According to TuxCare, KernelCare must not be run in parallel with Canonical Livepatch. Running both services simultaneously is not a meaningful compatibility test, but rather a disqualifying factor: First, the existing service must be removed according to the approved procedure, or the test platform must be disconnected. An internal comparison of the various procedures provides an overview of KernelCare, Ksplice, kpatch, and kGraft.

What a reliable test must demonstrate

A robust test begins with verifiable objectives rather than a blanket statement such as „Patch installed.“ Evidence must be provided of a supported, running kernel, an accessible and authorized patch source, and an up-to-date patch status. In addition, the team must record the effective security version reported by KernelCare. This evidence confirms the technical supply chain, but not yet the functionality of the application.

The second level of testing is the Application Health. Services must remain accessible, critical transactions must complete successfully, and interfaces must return the expected results. For database systems, replication and queries can be crucial; for web services, for example, authentication, background jobs, and external integrations should be included in the scope of testing.

For monitoring, it provides kcarectl --status Machine-readable exit codes. TuxCare assigns 0 to the latest patch level, 1 to no patches applied, 2 to new patches that have not yet been applied, and 3 to an unsupported kernel. These statuses are suitable for alert rules but must be evaluated in conjunction with kernel logs, service metrics, and technical reviews.

Also, distinguish between the booted version and the effective version. uname -r displays the booted kernel, while kcarectl --uname displays the secure kernel version reported by TuxCare. If this information is not properly accounted for in the scanner and CMDB, an effective live patch may appear as a missing update.

Approval requires complete technical documentation, successful application tests, and a representative load cycle. This can be a batch window, a typical peak load, or a scheduled failover. In the event of an unsupported kernel, an increasing number of errors, or failed functional tests, the rollout is halted and the issue is investigated; a positive agent status does not override such signals.

Set Up a Production-Ready Staging Baseline

A robust test begins with a staging host that replicates the intended target environment as closely as possible. Record the distribution, booted kernel, architecture, virtualization type, and enabled security mechanisms. The inventory should also include loaded or mission-critical kernel modules, storage and network paths, security and monitoring agents, and the core application components. Compatibility must always be verified for the kernel that is actually running, not just for the distribution.

Before making any changes, also document the application's current status: successful business transactions, error rates, response times, background jobs, and, if necessary, cluster membership or replication status. These Baseline makes it possible to trace subsequent discrepancies. Also check whether a backup or snapshot suitable for the application is available and how its restoration will be handled in practice; a VM snapshot does not replace a consistent database backup.

Close-up of a set-up staging workstation with a server and network cabling.
AI-generated stock image: A documented staging baseline provides reference values prior to the patch.

A lightweight test VM is useful for verifying the installation, registration, and availability of the patch source. However, it does not provide reliable insights into production-grade drivers, specialized modules, or load patterns. The Upstream Linux Livepatch framework technically classifies activations via a consistency transition; however, no specific KernelCare mechanism can be derived from this. Regardless, real-world workload profiles and additional operational components should be included in a representative staging test.

Test Objectives for the Staging Baseline and Their Limits of Detection
test objectiveEvidence in the test reportTypical detection limit
Detect the runtime environmentDocumentation on the kernel, architecture, virtualization, and relevant modulesDoes not yet indicate that a patch is available for this kernel build
Clarify recoverabilityBackup or snapshot procedures and responsibilities are documentedThe existence of a backup does not prove that the application was successfully restored
Check for technical patchabilityThe agent detects a supported kernel and can retrieve patch informationDoes not indicate whether the application is technically correct
Compare App HealthDefined transactions, metrics, and log checks before and after the patchCovers only the functions performed and the period observed
Monitor Load BehaviorTypical batch, peak load, or failover phase scheduledA brief idle test is no substitute for a load cycle

Do not set a blanket duration for the observation period. For a service with nightly imports, the test must include at least one such import; for a high-availability cluster, a controlled failover may be relevant. Define target values and termination criteria in advance. If new kernel messages, repeated agent errors, or business-related deviations occur, approval will not be granted, and the findings will be investigated before the next wave begins.

Correctly Evaluate Patch Status Using kcarectl

Record the system state before and after an approved patching operation using the same commands. This allows you to determine which kernel was booted, which agent version the host is using, and whether a patch set is actually active. The results—including the timestamp, host ID, and version of the application being tested—must be included in the change log or test log. A single success message from the installer is not sufficient proof.

The following queries are read-only and are suitable for taking an inventory. Execute them in the target environment using the permissions provided there. Only a later, deliberately planned update process changes the patch status; therefore, the output of these commands serves as a basis for comparison and monitoring, not the patching process itself.

Terminal
uname -r
kcarectl --version
kcarectl --info
kcarectl --patch-info
kcarectl --status
kcarectl --uname
The Significance of Key kcarectl Queries
CommandPurposeRelevant statementBorder
uname -rRecord the booted kernelDisplays the kernel release of the currently running systemDoes not display the security version achieved through Livepatch
kcarectl –versionTake Inventory of AgentsRecords the installed client versionDoes not indicate either support or an active patch status
kcarectl -infoGet Patch InformationDisplays information about the KernelCare statusDoes not replace a review of the application
kcarectl –patch-infoView patch detailsSupports the assignment of the patch setThere is no evidence of a technical function
kcarectl –statusCheck Machine-Readable StatusExit code 0 indicates the latest patch level; 1 indicates no patches; 2 indicates new patches that have not been applied; 3 indicates an unsupported kernelMust be evaluated in conjunction with agent and application monitoring
kcarectl –unameIssue an effective security releaseReturns the effective kernel version reported by TuxCareDoes not change the output of `uname -r`
kcarectl –checkSearch for a new patch setExit code 0 indicates that a new patch set is availableDoes not prove that the host has already been patched

It is particularly important to distinguish between booted and effective kernel version. A vulnerability scanner that only uname -r If assessed, this can give an outdated impression, even though a live patch provides the relevant fix. Therefore, align your inventory and compliance rules with the available TuxCare data, such as the effective version and the local CVE list at /proc/kcare/cvelist.

The following is suitable for alerts: kcarectl --status better than a simple text search in console output, because exit codes can be evaluated automatically. For example, a code 2 requires a decision on whether to roll out a new patch set within the specified window; code 3 indicates a compatibility or inventory issue. None of these codes replaces the need to review kernel logs, service metrics, and business transactions.

Stagger the rollout in a controlled manner across QA, Canary, and Production

A controlled rollout begins in a dedicated QA environment, then proceeds through a small, representative canary group, and is only expanded once stable results have been documented. Each wave undergoes the same status and application tests. The observation period depends on the load cycle: for batch systems, this involves a complete processing run; for clusters, it may include replication and a controlled failover.

During monitoring, you check error rates, latencies, kernel and agent messages, as well as quorum and replication, if applicable. Only after the release criteria have been met does the next group follow. The internal article explains additional fundamentals for deployment in live operations KernelCare Enterprise: Live Patching Without Downtime.

Rollout Options for KernelCare by Control Method and Application Area
OptionSuitable ApplicationsImportant Limitation
Standard Production FeedProduction Based on Our Own Approval ProcessContinued monitoring and staggered application are still required
Delayed Feed via PREFIXFixed delay of 12, 24, or 48 hoursThe delay stage is selected via the patch source
Test Feed via PREFIXDedicated QA or Canary systemsContains newer builds before the full testing process is complete
STICKY_PATCHLimit QA and production to a verified dateNot available for ePortal; key-based control not available for IP-based servers
STICKY_PATCHSET or UPDATE_DELAY starting with KernelCare 2.82Configure the patch set limit or a custom minimum ageAUTO options only work in Auto and Smart modes
ePortalCentralized control in controlled or isolated environmentsSetup, registration, accessibility, and policies remain prerequisites

Delayed feeds and UPDATE_DELAY solve similar problems at different levels. A feed is created via PREFIX Selected as a patch source with a fixed delay. UPDATE_DELAY In contrast, Patchsets holds back patches via the client configuration until they reach a specified minimum age. STICKY_PATCHSET Limits the client to a specific maximum patch set version.

A manual kcarectl --update Downloads the latest patch set and applies it to the running kernel. Use this command only on approved test systems or during a designated maintenance window. Be sure to back up the baseline values beforehand and perform technical and functional tests immediately afterward.

ePortal can centrally manage patch sets and deployment. According to TuxCare, when automatic updates are enabled, clients check for available patch sets every four hours. This does not guarantee an execution time: availability, registration, policies, and kernel compatibility must be monitored for each wave.

For each wave, record the patch status, selected hosts, monitoring window, test results, and the person responsible for approval. If any discrepancies are found, the rollout is paused. These Canary Release limits the scope of unexpected effects, but does not replace either the compatibility check or the scheduled reboot cycle.

Testing Secure Boot and Critical Special Cases

Server with Secure Boot belong in a separate test group. The agent requires a functioning chain of trust for its kernel modules; a successful installation run does not yet prove this. TuxCare specifies Agent version 3.0-2 as the minimum version required for the automated Secure Boot process on supported RPM systems. This specification does not apply as a general minimum version for KernelCare, nor does it apply to manual MOK registration.

For the automated method, EFI boot, shim, and enabled Secure Boot must be present, among other requirements. According to TuxCare, this process is not intended for Debian and Ubuntu. Therefore, record the distribution, boot mode, and agent version before the test, and do not treat a different platform as merely a configuration variant, but rather as a separate path that must be evaluated manually.

The test does not end until after a scheduled restart. Then use the tool described by TuxCare to verify mokutil or by checking appropriate kernel logs to determine whether the certificate is actually available in the chain of trust. Only then is a controlled LivePatch retrieval performed on this host, using the same functional and technical checks as in the rest of the QA wave.

An administrator checks the hardware and cabling during a Secure Boot maintenance check.
AI-generated illustrative image: Secure-boot systems require separate validation with a scheduled reboot.

Systems with proprietary drivers, storage or network modules, eBPF programs, security software, and monitoring agents also require their own representative test coverage. This is not a general statement about incompatibility. From a technical standpoint, the upstream Linux Livepatch framework describes consistency transitions for affected tasks; however, this does not prove that KernelCare uses the same mechanism on every supported platform.

Therefore, simulate the combinations that actually occur in production: for example, multipath storage under load, encrypted network connections, security agents, and the failover role of a cluster node. Document loaded modules, kernel messages, and the application and cluster status before and after applying the patch. A stripped-down test VM without these components can confirm the agent installation but cannot provide reliable insights into this class of systems.

Monitoring, Error Analysis, and Secure Escalation

Monitor live patching on two levels: The machine-readable Patch Status shows the agent's status, while kernel logs, error rates, latencies, and cluster status reflect the application's operation. An up-to-date patch status does not rule out the possibility of a concurrent application failure or business-related deviation. Therefore, alerting and approval processes must integrate both levels and investigate the cause of a deviation separately.

For automated triage, it provides kcarectl --status Defined exit codes: 0 indicates the latest patch level, 1 indicates no patches applied, 2 indicates available but not yet applied patches, and 3 indicates an unsupported kernel. Code 3 requires a compatibility check first; Code 2 is not an application error, but must be evaluated against the planned rollout and update policy.

When discrepancies occur, first collect data that can be correlated in time: status and patch information, agent reports, kernel logs, the time of the query, affected workloads, and changes to modules or infrastructure. For cluster nodes, this includes membership, replication status, and failover events. This data distinguishes a patch status from a concurrent application or network failure and makes a support case traceable.

TuxCare documents kcarectl --force as an option along with an update that forces the application of a patch if some threads cannot be frozen. The upstream Linux documentation warns of potential damage when using its own force mechanism, requires a scheduled reboot afterward, and advises against further live patches. However, it does not provide evidence that kcarectl --force uses the same semantics internally. Therefore, the product-specific TuxCare support instructions and the diagnosis of the specific host are decisive; this option is not suitable as a standard rollout or troubleshooting measure.

Plan a reboot strategy and documented approval

Live patching reduces the time it takes to address supported kernel vulnerabilities, but does not modify the installed kernel package. New kernel packages, hardware support, driver or firmware changes, and functional kernel improvements still require regular package management and scheduled reboots.

Therefore, define a reboot schedule for each platform class. KernelCare provides patches for a specific kernel only as long as its manufacturer continues to provide security updates for that series. A maintenance window also ensures that the booted kernel, the loaded drivers, and the documented target state are brought back into alignment.

TuxCare documents kcarectl --unload for unloading KernelCare patches. This does not imply a general guarantee of complete recovery. The upstream documentation indicates that, for Atomic Replace and cumulative live patches, state changes can make rollback difficult; however, it does not automatically describe the specific implementation of every KernelCare version.

Before performing an uninstallation, therefore, check the documentation for the installed agent version and, if necessary, coordinate troubleshooting steps with TuxCare. The reliable Return Point What remains is a defined, tested boot kernel with a scheduled reboot and, if necessary, a consistency check or application recovery.

The approval of a rollout wave documents the supported kernel, the patch status, application tests performed, relevant load cycles, logs, responsible parties, and cancellation criteria. It does not constitute a blanket commitment to future patch sets. Changes to the kernel, modules, or application may require repeat QA and canary testing.

  • Document the supported kernel, patch source, and current patch version.
  • Verify application checks, load cycles, kernel logs, and cluster status to ensure there are no unexplained discrepancies.
  • Define the rollout phase, responsible parties, alerting procedures, and termination criteria.
  • Schedule the next kernel update, including the maintenance window, boot kernel, and restart check.

This makes the operational decision clear: A successful live patch allows for the controlled continuation of the respective wave. Unresolved technical or operational issues, on the other hand, lead to a hold, analysis, or a planned restart. Reboot planning is part of the security and recovery strategy; it is not an admission that a live patch has failed.

Sources and Current State of Knowledge

Status of the research:

Research status: September 28, 2026. Before use, verify information regarding support, agent versions, feeds, and commands against the current TuxCare documentation and the kernel actually in use.

https://docs.tuxcare.com/live-patching-services/

https://docs.kernel.org/6.12/livepatch/livepatch.html

https://docs.tuxcare.com/eportal/

https://docs.kernel.org/6.0/livepatch/cumulative-patches.html

Current articles

Administrator in a hosting operations room, in addition to server technology
Servers and Virtual Machines

CloudLinux OS 9: Features and Limitations in Shared Hosting

CloudLinux OS 9 modernizes the system foundation for shared hosting. However, the license, edition, installed components, and control panel integration remain critical—especially for LVE, CageFS, Isolates, and Shared Pro.