Understanding V4L2 Buffer Formats-Free Linux Device Drivers Course

PREV_LEC  |  NEXT_LEC

Understanding V4L2 Buffer Formats

Free Linux Kernel Development Course — V4L2 User Space API, Part 3

This lecture is part of Embedded Pathashala’s free Linux kernel development course and continues our free Linux device drivers course series on the V4L2 user space API. If you are looking for a free embedded Linux course that explains real ioctl-level mechanics instead of just theory, this is where V4L2 buffer format management stops being confusing. In the last two lectures we opened a video device node and queried its capabilities. Now we tackle the two things every capture application must get right before a single frame ever arrives: how buffers move between kernel and user space, and how the video format itself is negotiated with VIDIOC_G_FMT and VIDIOC_S_FMT.

Topics Covered

V4L2 buffer queues struct v4l2_format VIDIOC_G_FMT / VIDIOC_S_FMT pixel format negotiation multiplanar V4L2 free linux kernel development course

What You Will Learn

  • How the V4L2 driver-side and user-side buffer queues actually move a buffer from empty to filled
  • The full modern struct v4l2_format layout, including buffer types added since this API was first documented
  • How to query a device’s current format with VIDIOC_G_FMT and change it with VIDIOC_S_FMT
  • Why the driver is allowed to silently adjust the format you asked for, and how to detect that
  • How to inspect and change the capture frame rate with VIDIOC_G_PARM / VIDIOC_S_PARM
  • A working original C demo, ep_v4l2fmt, that negotiates a format against a real capture device

Prerequisites

  • Completion of the previous lectures on the V4L2 device open sequence and VIDIOC_QUERYCAP
  • A Linux system (kernel 6.x recommended) with a V4L2 capture device — a USB webcam is enough
  • Basic C and familiarity with ioctl() calls on file descriptors
  • v4l-utils installed for cross-checking (sudo apt install v4l-utils)

The Two-Queue Model Behind Every V4L2 Buffer

Before any format negotiation makes sense, you need a clear mental model of how a V4L2 buffer physically travels through the system. This is the foundation of V4L2 buffer format management, and skipping it is the single biggest reason beginners get confused later when we cover VIDIOC_QBUF and VIDIOC_DQBUF in depth.

Every streaming V4L2 device keeps two logical buffer queues. The driver owns an incoming queue — buffers that are empty and waiting to be filled with captured data. The application owns an outgoing queue — buffers the driver has already filled and handed back for processing. A buffer never sits in both places at once; it moves in a strict cycle:

V4L2 Buffer Lifecycle
[ App allocates buffer ]  →  [ Driver’s incoming queue ]  →  [ Hardware fills buffer ]  →  [ App’s outgoing queue ]  →  [ App reads/copies frame ]  →  back to incoming queue

An application enqueues an empty buffer into the driver’s queue with VIDIOC_QBUF. The driver fills buffers strictly in the order it received them — there is no reordering. Once a buffer is full, the driver silently moves it out of its own queue and into the application’s queue. When the application calls VIDIOC_DQBUF, the kernel looks for a ready buffer in that outgoing queue: if one exists it is handed back immediately, otherwise the call blocks (or returns EAGAIN in non-blocking mode) until a frame is ready. After the application finishes using the frame, it must call VIDIOC_QBUF again to return the same buffer to the driver — otherwise the driver eventually runs out of buffers to fill and frames get dropped at the hardware level.

The full sequence an application follows, at a high level, is: negotiate the format, call VIDIOC_REQBUFS to tell the driver how many buffers you want, allocate/map those buffers, queue every one of them with VIDIOC_QBUF, then call VIDIOC_STREAMON to start the hardware. We cover REQBUFS, memory models, and the full queue/dequeue loop in detail in upcoming lectures — this lecture focuses on the piece that has to happen first: format negotiation.

struct v4l2_format on a Modern Kernel

Every ioctl that touches buffer geometry — cropping, requesting buffers, queueing, dequeueing, or querying format — needs to know which kind of stream it is operating on. That is exactly what struct v4l2_format exists for. Its shape has not changed structurally in years, though the set of valid type values has grown as V4L2 gained support for more capture styles:

struct v4l2_format {
    __u32 type;
    union {
        struct v4l2_pix_format          pix;      /* V4L2_BUF_TYPE_VIDEO_CAPTURE */
        struct v4l2_pix_format_mplane   pix_mp;   /* V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE */
        struct v4l2_window              win;      /* V4L2_BUF_TYPE_VIDEO_OVERLAY */
        struct v4l2_vbi_format          vbi;      /* V4L2_BUF_TYPE_VBI_CAPTURE */
        struct v4l2_sliced_vbi_format   sliced;   /* V4L2_BUF_TYPE_SLICED_VBI_CAPTURE */
        struct v4l2_sdr_format          sdr;      /* V4L2_BUF_TYPE_SDR_CAPTURE */
        struct v4l2_meta_format         meta;     /* V4L2_BUF_TYPE_META_CAPTURE */
        __u8                            raw_data[200];
    } fmt;
};

The type field is set by the application and tells the kernel which member of the union is valid. Almost every modern sensor and ISP driver you will run into today prefers the multiplanar types — V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE using pix_mp — because it lets a single buffer format describe formats where luma and chroma live in separate memory planes, such as NV12M or YUV420M. The classic single-plane V4L2_BUF_TYPE_VIDEO_CAPTURE with pix is still fully supported and remains the simplest starting point, which is what we use in this lecture’s demo. The complete, current list of buffer types lives in the enum v4l2_buf_type definition inside the kernel’s videodev2.h uAPI header — it is worth reading once so you recognize every type when you see it in driver source.

Important note: A modern kernel also exposes V4L2_BUF_TYPE_META_CAPTURE and V4L2_BUF_TYPE_META_OUTPUT for sensor metadata streams (exposure, gain, embedded statistics) that ship alongside the pixel stream on ISP-based platforms — these did not exist when V4L2 buffer negotiation was first documented in older texts, so don’t be surprised to see them in current driver code.

Querying the Current Format with VIDIOC_G_FMT

Applications almost always query the driver’s current format before changing anything, rather than guessing blindly. The rule is simple: pass a freshly zeroed struct v4l2_format with only type filled in, and the kernel fills in the rest.

struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));

fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;

if (ioctl(fd, VIDIOC_G_FMT, &fmt) == -1) {
    perror("VIDIOC_G_FMT");
    close(fd);
    exit(EXIT_FAILURE);
}

printf("Current format: %ux%u, fourcc=%c%c%c%c\n",
       fmt.fmt.pix.width, fmt.fmt.pix.height,
       fmt.fmt.pix.pixelformat & 0xFF,
       (fmt.fmt.pix.pixelformat >> 8) & 0xFF,
       (fmt.fmt.pix.pixelformat >> 16) & 0xFF,
       (fmt.fmt.pix.pixelformat >> 24) & 0xFF);

Zeroing the structure matters more than it looks. Any garbage left in the union from the stack can be misread by the driver on some code paths, and it makes debugging painful when fields you never touched appear to have random values.

Common V4L2 Pixel Formats

Once you have the current format, you typically only change the fields you care about and send the whole structure back. V4L2 fourcc pixel formats you will run into constantly on real hardware include:

FourccDescriptionTypical use
V4L2_PIX_FMT_YUYVYUV 4:2:2, packed/interleavedMost USB UVC webcams
V4L2_PIX_FMT_UYVYYUV 4:2:2, byte-swapped packedAnalog/BT.656 style capture bridges
V4L2_PIX_FMT_NV12YUV 4:2:0, semi-planarHardware video encoders/ISPs
V4L2_PIX_FMT_NV16YUV 4:2:2, semi-planarISP intermediate formats
V4L2_PIX_FMT_RGB24RGB 8:8:8, packedSimple display pipelines, testing
V4L2_PIX_FMT_SRGGB1010-bit raw BayerRaw sensor output before ISP demosaic

Setting a New Format with VIDIOC_S_FMT

To change the format, fill in the fields you need and issue VIDIOC_S_FMT. Because some kernels still route this through a compatibility wrapper on certain drivers, using the xioctl()-style EINTR-safe wrapper from our previous lecture is good practice here too.

#define CAP_WIDTH   1280
#define CAP_HEIGHT  720
#define CAP_PIXFMT  V4L2_PIX_FMT_YUYV

struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));

fmt.type                = V4L2_BUF_TYPE_VIDEO_CAPTURE;
fmt.fmt.pix.width       = CAP_WIDTH;
fmt.fmt.pix.height      = CAP_HEIGHT;
fmt.fmt.pix.pixelformat = CAP_PIXFMT;
fmt.fmt.pix.field       = V4L2_FIELD_NONE;

if (ep_xioctl(fd, VIDIOC_S_FMT, &fmt) == -1) {
    perror("VIDIOC_S_FMT");
    close(fd);
    exit(EXIT_FAILURE);
}

Notice we did not fill in bytesperline, sizeimage, or colorspace — for VIDIOC_S_FMT on most modern capture drivers, leaving these at zero tells the driver to calculate them for you based on width, height, and pixel format. This is safer than computing them by hand, since only the driver knows the true stride requirements of the underlying hardware.

Important note: A successful return from VIDIOC_S_FMT does not mean your exact request was honored. A device may not support every width/height/pixel-format combination. When that happens, the driver silently substitutes the closest values it does support and writes them back into the same fmt structure. Always re-check the structure after the call.

if (fmt.fmt.pix.pixelformat != CAP_PIXFMT)
    fprintf(stderr, "Driver rejected requested pixel format, using its own choice instead\n");

if (fmt.fmt.pix.width != CAP_WIDTH || fmt.fmt.pix.height != CAP_HEIGHT)
    fprintf(stderr, "Driver adjusted resolution to %ux%u\n",
            fmt.fmt.pix.width, fmt.fmt.pix.height);

Controlling Frame Rate with VIDIOC_G_PARM / VIDIOC_S_PARM

Format negotiation covers geometry and pixel layout, but not timing. Frame rate is a separate negotiation using struct v4l2_streamparm:

struct v4l2_streamparm parm;
memset(&parm, 0, sizeof(parm));
parm.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;

if (ioctl(fd, VIDIOC_G_PARM, &parm) == -1) {
    perror("VIDIOC_G_PARM");
} else if (parm.parm.capture.capability & V4L2_CAP_TIMEPERFRAME) {
    parm.parm.capture.timeperframe.numerator   = 1;
    parm.parm.capture.timeperframe.denominator = 30; /* request 30 fps */

    if (ioctl(fd, VIDIOC_S_PARM, &parm) == -1)
        perror("VIDIOC_S_PARM");
}

Check V4L2_CAP_TIMEPERFRAME before attempting VIDIOC_S_PARM — many simple UVC devices don’t support arbitrary frame-rate control at all, and calling it blindly just wastes a round trip. You can enumerate exactly which frame intervals a device supports for a given resolution and pixel format with VIDIOC_ENUM_FRAMEINTERVALS, which we will use directly in the v4l2-ctl lecture later in this chapter.

Demo: ep_v4l2fmt — Original Format Negotiation Tool

Here is a complete, original demo application, ep_v4l2fmt, that opens a capture device, prints its current format, requests 1280×720 YUYV, and reports whether the driver honored the request.

/* ep_v4l2fmt.c — EmbeddedPathashala original demo, not derived from any book source */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/ioctl.h>
#include <linux/videodev2.h>

static int ep_xioctl(int fd, unsigned long req, void *arg)
{
    int r;
    do {
        r = ioctl(fd, req, arg);
    } while (r == -1 && errno == EINTR);
    return r;
}

int main(int argc, char *argv[])
{
    const char *dev = (argc > 1) ? argv[1] : "/dev/video0";
    int fd = open(dev, O_RDWR);
    if (fd == -1) {
        perror("open");
        return 1;
    }

    struct v4l2_format fmt;
    memset(&fmt, 0, sizeof(fmt));
    fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;

    if (ep_xioctl(fd, VIDIOC_G_FMT, &fmt) == -1) {
        perror("VIDIOC_G_FMT");
        close(fd);
        return 1;
    }

    printf("[%s] Current format: %ux%u fourcc=%c%c%c%c\n", dev,
           fmt.fmt.pix.width, fmt.fmt.pix.height,
           fmt.fmt.pix.pixelformat & 0xFF,
           (fmt.fmt.pix.pixelformat >> 8) & 0xFF,
           (fmt.fmt.pix.pixelformat >> 16) & 0xFF,
           (fmt.fmt.pix.pixelformat >> 24) & 0xFF);

    fmt.fmt.pix.width       = 1280;
    fmt.fmt.pix.height      = 720;
    fmt.fmt.pix.pixelformat = V4L2_PIX_FMT_YUYV;
    fmt.fmt.pix.field       = V4L2_FIELD_NONE;

    if (ep_xioctl(fd, VIDIOC_S_FMT, &fmt) == -1) {
        perror("VIDIOC_S_FMT");
        close(fd);
        return 1;
    }

    printf("[%s] Granted format:  %ux%u fourcc=%c%c%c%c\n", dev,
           fmt.fmt.pix.width, fmt.fmt.pix.height,
           fmt.fmt.pix.pixelformat & 0xFF,
           (fmt.fmt.pix.pixelformat >> 8) & 0xFF,
           (fmt.fmt.pix.pixelformat >> 16) & 0xFF,
           (fmt.fmt.pix.pixelformat >> 24) & 0xFF);

    if (fmt.fmt.pix.pixelformat != V4L2_PIX_FMT_YUYV)
        fprintf(stderr, "Note: driver substituted a different pixel format\n");

    close(fd);
    return 0;
}

Build and run it against a real capture device:

$ gcc -Wall -o ep_v4l2fmt ep_v4l2fmt.c
$ ./ep_v4l2fmt /dev/video0
[/dev/video0] Current format: 640x480 fourcc=YUYV
[/dev/video0] Granted format:  1280x720 fourcc=YUYV

If your webcam does not support 1280×720 at YUYV, you will see the granted resolution snap to the closest supported mode instead — that is expected driver behavior, not a bug in the demo. You can cross-check the result independently:

$ v4l2-ctl -d /dev/video0 --get-fmt-video
Format Video Capture:
	Width/Height      : 1280/720
	Pixel Format      : 'YUYV' (YUYV 4:2:2)
	Field             : None

Common Mistakes and Troubleshooting

  • Not zeroing the structure: passing a stack-garbage struct v4l2_format into VIDIOC_S_FMT can leave unrelated union fields set to unpredictable values.
  • Assuming the request was honored: always re-read width, height, and pixelformat from the structure after VIDIOC_S_FMT returns.
  • Calling S_FMT after STREAMON: format must be negotiated before VIDIOC_REQBUFS/VIDIOC_STREAMON — most drivers return EBUSY if you try to change format mid-stream.
  • Ignoring V4L2_CAP_TIMEPERFRAME: calling VIDIOC_S_PARM on a device that doesn’t support frame-rate control wastes a call and can be mistaken for a driver bug.
  • Forgetting multiplanar devices: on an MPLANE-only sensor driver, fmt.fmt.pix is the wrong union member — you must use fmt.fmt.pix_mp and the type V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE, or every ioctl will fail with EINVAL.

Best Practices

  • Always query with VIDIOC_G_FMT before setting, so you know the driver’s defaults and only change what you actually need.
  • Validate the granted format before proceeding to VIDIOC_REQBUFS — don’t assume your requested resolution stuck.
  • Prefer letting the driver compute bytesperline/sizeimage rather than hand-calculating them, unless you have a specific hardware alignment requirement.
  • Check device capabilities (from the earlier VIDIOC_QUERYCAP lecture) for V4L2_CAP_VIDEO_CAPTURE_MPLANE before assuming single-plane pix is correct.
  • Wrap every ioctl call in an EINTR-safe retry helper — signal interruption is common in real capture applications.

Interview Questions

Why does V4L2 use two separate buffer queues instead of one?

Separating the driver’s incoming (empty) queue from the application’s outgoing (filled) queue lets the kernel and hardware fill buffers asynchronously in strict FIFO order while the application processes previously filled buffers at its own pace, without either side blocking the other unnecessarily.

What happens if VIDIOC_S_FMT is called with an unsupported resolution?

The ioctl still returns success. The driver adjusts the requested width, height, or pixel format to the closest combination it actually supports and writes those adjusted values back into the same struct v4l2_format passed by the caller — it is the application’s responsibility to check the returned values.

How do you know whether to use struct v4l2_pix_format or v4l2_pix_format_mplane?

Check the device’s reported capabilities from VIDIOC_QUERYCAP. If V4L2_CAP_VIDEO_CAPTURE_MPLANE is set, the device expects the multiplanar pix_mp member with buffer type V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE; using the single-plane pix member against such a device returns EINVAL.

Summary and Key Takeaways

V4L2 buffer format management rests on two ideas that every capture application must respect: buffers move through a strict two-queue cycle between driver and application, and format negotiation via VIDIOC_G_FMT/VIDIOC_S_FMT is a request, not a guarantee — the driver always has the final say on what it can actually deliver. Getting this right before touching VIDIOC_REQBUFS or streaming ioctls saves hours of confusing debugging later. In the next lecture in this free Linux kernel development course, we extend this into VIDIOC_G_PARM/VIDIOC_S_PARM in more depth and start allocating real buffers with VIDIOC_REQBUFS.

Frequently Asked Questions

What is the difference between VIDIOC_G_FMT and VIDIOC_S_FMT?

VIDIOC_G_FMT queries the device’s current format without changing anything. VIDIOC_S_FMT requests a new format; the driver may adjust the requested values to the nearest supported configuration.

Do I need to call VIDIOC_G_FMT before VIDIOC_S_FMT?

It is not strictly required, but it is good practice — it lets you start from the driver’s defaults and change only the fields you care about instead of guessing every field from scratch.

Why does the driver change my requested width and height?

Hardware capture pipelines only support a fixed set of resolutions and formats. When your request doesn’t exactly match one, the driver rounds to the closest supported combination and reports the actual values back to you.

What is V4L2_FIELD_NONE and why set it?

It tells the driver the stream is fully progressive with no interlacing, which is the correct setting for essentially all modern USB and MIPI CSI-2 camera sensors.

Can I change the format while streaming is active?

No. Format must be set before VIDIOC_REQBUFS and VIDIOC_STREAMON. Most drivers return EBUSY if you attempt VIDIOC_S_FMT once streaming has started.

What does V4L2_CAP_TIMEPERFRAME mean?

It is a capability flag reported via VIDIOC_G_PARM that tells you whether the driver supports changing the capture frame rate at all. If it’s not set, VIDIOC_S_PARM has no effect.

What is the difference between V4L2_PIX_FMT_NV12 and V4L2_PIX_FMT_YUYV?

YUYV is a packed 4:2:2 format with luma and chroma interleaved in one plane. NV12 is a semi-planar 4:2:0 format with a separate luma plane and an interleaved chroma plane — common on hardware ISPs and video encoders.

Is this course free?

Yes. This lecture is part of Embedded Pathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded Linux course material.

Continue Your Free Linux Kernel Development Course

Master the complete V4L2 user space API step by step, with original demos and real kernel behavior — no vendor code, no fluff.

Next Lecture Browse Full Course Index

PREV_LEC  |  NEXT_LEC

Leave a Reply

Your email address will not be published. Required fields are marked *