How Does a Slab Constructor Work in Linux? – Free Linux Device Drivers Training Online

 

Custom Slab Caches & Constructor Functions in the Linux Kernel
Free Linux Kernel Development Course — Memory Management Module

If you are building a free Linux kernel development course mental checklist, custom slab caches belong near the top of the memory management section. Almost every serious Linux driver that allocates the same kind of object thousands of times per second — network buffers, filesystem inodes, block I/O requests — avoids plain kmalloc() and instead creates its own dedicated slab cache. In this lesson of our free Linux kernel development course we build one from scratch, attach a constructor function to it, and verify exactly how much memory the kernel really hands out per object on a modern kernel.

slab allocator
kmem_cache_create
SLUB
kernel object caching
Linux device driver memory
free embedded systems course

What You Will Learn

  • Why drivers create a custom slab cache instead of using kmalloc()
  • How to allocate a cache with kmem_cache_create() on a current kernel
  • What a constructor function does and when the kernel actually calls it
  • Why the size the kernel reports for your object is not the size it actually reserves
  • How to inspect live slab cache statistics using slabtop and /proc/slabinfo on modern kernels
  • Common mistakes beginners make when working with custom kernel caches

Prerequisites

Before this lesson, you should already be comfortable with:

  • Writing and loading a basic “Hello World” loadable kernel module
  • module_param() and reading kernel log output with dmesg
  • Basic C structures, pointers, and the difference between stack and heap memory

If any of these feel shaky, go back one lecture in this free Linux kernel development course before continuing.

Why a Custom Slab Cache Instead of kmalloc()

kmalloc() is convenient, but it hands you memory from general-purpose size buckets (32, 64, 128 bytes, and so on). If your driver repeatedly allocates and frees an object of an odd size, you end up wasting the gap between your object size and the nearest bucket, and the allocator gets no chance to specialize for your access pattern. A custom slab cache solves this: the kernel carves out memory specifically shaped for your structure, keeps a pool of ready-to-use objects, and — this is the part most tutorials skip — lets you run initialization code automatically every time a fresh object is handed to that pool.

The Slab Allocator Landscape on a Modern Kernel

Older material on this topic usually mentions three competing slab implementations. On a current kernel you really only need to know about one of them:

Allocator Status on current kernels
SLOB Removed. It targeted very small embedded systems and was dropped from the tree.
SLAB Removed. The original allocator was retired in favor of SLUB, which had matched or beaten it on most workloads for years.
SLUB The only slab allocator in the current kernel. Every example in this lesson assumes SLUB.

Why this matters practically: some very old blog posts tell you to read cache statistics with vmstat -m. That command reported meaningful numbers under the old SLAB allocator; on a SLUB-only kernel you should instead reach for slabtop and /proc/slabinfo, which we use later in this lesson.

Step 1: Defining the Object Structure

Let’s build a small demo module that caches a fictitious “network buffer descriptor” object. First, the structure we want cached:

struct netbuf_ctx {
    u32  seq_num;
    u16  len;
    u8   flags;
    char origin[32];
};

static struct kmem_cache *netbuf_cachep;

Step 2: Writing a Constructor Function

A constructor is a callback you hand to the slab layer. The kernel invokes it once, automatically, whenever it physically carves out a brand-new object for your cache — not every time you call the allocation function. This distinction trips up almost every beginner, so we’ll verify it with our own log output shortly.

static void netbuf_ctor(void *obj)
{
    struct netbuf_ctx *nb = obj;

    memset(nb, 0, sizeof(*nb));
    strscpy(nb->origin, "netbuf-cache", sizeof(nb->origin));
    pr_info("netbuf_ctor: new backing object initialised at %px\n", nb);
}

Step 3: Creating the Cache

static int __init netbuf_demo_init(void)
{
    netbuf_cachep = kmem_cache_create(
        "netbuf_ctx_cache",
        sizeof(struct netbuf_ctx),
        0,                       /* natural alignment */
        SLAB_HWCACHE_ALIGN,
        netbuf_ctor);

    if (!netbuf_cachep)
        return -ENOMEM;

    pr_info("netbuf demo: cache created, reported object size = %u bytes\n",
            kmem_cache_size(netbuf_cachep));

    return 0;
}
How an Allocation Request Flows Through the Cache
Driver calls
kmem_cache_alloc()
→
Free object
already in pool?
→
No → SLUB carves
new memory,
runs constructor
→
Object handed
to driver

Notice the “Free object already in pool?” step. If SLUB already has a spare object sitting around from an earlier free, it hands that one straight back to you — the constructor is skipped because the object was already initialised the first time it was carved out. This is exactly why you should treat a constructor as one-time setup for the memory shape, never as per-allocation business logic.

Step 4: Allocating and Freeing an Object

struct netbuf_ctx *nb = kmem_cache_alloc(netbuf_cachep, GFP_KERNEL);
if (!nb)
    return -ENOMEM;

pr_info("netbuf demo: origin field reads back as '%s'\n", nb->origin);

kmem_cache_free(netbuf_cachep, nb);

And in your module’s exit path, always destroy the cache:

kmem_cache_destroy(netbuf_cachep);

Step 5: The Reported Size Is Not the Real Size

Here is the detail that catches almost everyone off guard. kmem_cache_size() and sizeof() tell you the logical object size — what your structure needs. SLUB, however, rounds objects up to fit its internal slab geometry, adds debugging/red-zone padding when certain kernel configs are enabled, and may share a common cache with other similarly sized objects entirely. The memory actually reserved per object is almost always larger than what sizeof() reports.

On a modern kernel, do not reach for vmstat -m to check this — it belongs to the retired SLAB era. Use one of these instead:

$ sudo slabtop -o | grep netbuf_ctx_cache

$ cat /proc/slabinfo | grep netbuf_ctx_cache

Both commands show you the active object count, the number of objects per slab page, and the real per-object size SLUB is using — which will typically be noticeably larger than your raw sizeof(struct netbuf_ctx) value once alignment and any active slab debugging options are factored in.

Common Mistakes and Troubleshooting

Mistake Why it hurts you
Putting per-request logic inside the constructor It only runs when a new object is physically created, not on every kmem_cache_alloc() call, so recycled objects silently skip it.
Forgetting kmem_cache_destroy() on module unload Leaves a dangling cache reference; reloading the module can trigger warnings or leaks.
Assuming sizeof() equals real memory used Leads to wrong capacity planning under memory pressure; always confirm with slabtop.
Using vmstat -m on a modern kernel That view was tied to the retired SLAB allocator and no longer reflects SLUB accurately.

Best Practices

  • Only create a custom cache when you allocate the same structure very frequently — otherwise kmalloc() is simpler and fine.
  • Keep constructors free of anything that depends on request-specific data.
  • Always pair kmem_cache_create() with a corresponding kmem_cache_destroy() in your exit path.
  • Use SLAB_HWCACHE_ALIGN for objects that are touched heavily on hot paths, to reduce false sharing across CPU cores.

Performance Considerations

Custom caches shine under high allocation churn — think packet buffers or per-connection state in a driver handling thousands of events per second. For occasional, low-frequency allocations, the bookkeeping overhead of a dedicated cache is not worth it; plain kmalloc() wins there.

Security Considerations

Never assume a freshly allocated object is zeroed unless your constructor explicitly zeroes it (as we did with memset() above) or you request a zeroing allocation flag. Stale data left over from a previous user of that memory slot is a classic source of kernel information-leak bugs.

Summary / Key Takeaways

  • Custom slab caches exist to specialize memory allocation for one frequently-used structure.
  • A constructor runs once per physically created object, not once per allocation call.
  • SLUB is the only slab allocator on current kernels — SLAB and SLOB are gone.
  • Use slabtop / /proc/slabinfo, not vmstat -m, to inspect real per-object memory usage today.

Frequently Asked Questions

Q1. Does the constructor run every time I call kmem_cache_alloc()?
No. It only runs when SLUB has to carve a genuinely new object out of a slab page. If a previously freed object is reused, the constructor is skipped.

Q2. Is SLAB still available if I set a kernel config option?
No, once an allocator is removed from the kernel tree it cannot be re-enabled through configuration; you would need to run an older kernel version entirely.

Q3. Why does slabtop show a bigger size than my sizeof(struct)?
SLUB rounds up for alignment and may add extra bytes for debugging metadata when certain kernel debug options are active.

Q4. Can I pass arguments into a constructor function?
No, the constructor signature only receives the object pointer. Anything request-specific must be set after allocation, not inside the constructor.

Q5. Is creating a custom slab cache ever a bad idea?
Yes, for objects allocated rarely. The extra bookkeeping isn’t worth it unless allocation frequency is genuinely high.

Q6. What replaced vmstat -m for checking slab memory today?
Use slabtop for a live view or read /proc/slabinfo directly for scripting.

Q7. Do I need root privileges to inspect slab caches?
Typically yes — both slabtop and reading detailed /proc/slabinfo fields usually require root or elevated privileges on most distributions.

Conclusion

Custom slab caches are one of those Linux kernel internals topics that looks intimidating in the reference manual but becomes very approachable once you write and run one yourself. You now know how to create a cache, attach a constructor that behaves correctly under object reuse, and — critically for a modern kernel — how to verify actual memory usage with the right tools instead of retired ones. In the next lecture of this free Linux kernel development course, we take this a step further and register a slab shrinker so your cache cooperates properly under system memory pressure.

Continue the Free Linux Kernel Development Course

More free lessons on Linux kernel programming, device drivers, and embedded systems at EmbeddedPathashala.

Explore All Free Courses

 

1 Comment

Leave a Reply

Your email address will not be published. Required fields are marked *