Understanding V4L2 Buffer Formats
Free Linux Kernel Development Course — V4L2 User Space API, Part 3
This lecture is part of Embedded Pathashala’s free Linux kernel development course and continues our free Linux device drivers course series on the V4L2 user space API. If you are looking for a free embedded Linux course that explains real ioctl-level mechanics instead of just theory, this is where V4L2 buffer format management stops being confusing. In the last two lectures we opened a video device node and queried its capabilities. Now we tackle the two things every capture application must get right before a single frame ever arrives: how buffers move between kernel and user space, and how the video format itself is negotiated with VIDIOC_G_FMT and VIDIOC_S_FMT.
Topics Covered
What You Will Learn
- How the V4L2 driver-side and user-side buffer queues actually move a buffer from empty to filled
- The full modern
struct v4l2_formatlayout, including buffer types added since this API was first documented - How to query a device’s current format with
VIDIOC_G_FMTand change it withVIDIOC_S_FMT - Why the driver is allowed to silently adjust the format you asked for, and how to detect that
- How to inspect and change the capture frame rate with
VIDIOC_G_PARM/VIDIOC_S_PARM - A working original C demo,
ep_v4l2fmt, that negotiates a format against a real capture device
Prerequisites
- Completion of the previous lectures on the V4L2 device open sequence and
VIDIOC_QUERYCAP - A Linux system (kernel 6.x recommended) with a V4L2 capture device — a USB webcam is enough
- Basic C and familiarity with
ioctl()calls on file descriptors v4l-utilsinstalled for cross-checking (sudo apt install v4l-utils)
The Two-Queue Model Behind Every V4L2 Buffer
Before any format negotiation makes sense, you need a clear mental model of how a V4L2 buffer physically travels through the system. This is the foundation of V4L2 buffer format management, and skipping it is the single biggest reason beginners get confused later when we cover VIDIOC_QBUF and VIDIOC_DQBUF in depth.
Every streaming V4L2 device keeps two logical buffer queues. The driver owns an incoming queue — buffers that are empty and waiting to be filled with captured data. The application owns an outgoing queue — buffers the driver has already filled and handed back for processing. A buffer never sits in both places at once; it moves in a strict cycle:
An application enqueues an empty buffer into the driver’s queue with VIDIOC_QBUF. The driver fills buffers strictly in the order it received them — there is no reordering. Once a buffer is full, the driver silently moves it out of its own queue and into the application’s queue. When the application calls VIDIOC_DQBUF, the kernel looks for a ready buffer in that outgoing queue: if one exists it is handed back immediately, otherwise the call blocks (or returns EAGAIN in non-blocking mode) until a frame is ready. After the application finishes using the frame, it must call VIDIOC_QBUF again to return the same buffer to the driver — otherwise the driver eventually runs out of buffers to fill and frames get dropped at the hardware level.
The full sequence an application follows, at a high level, is: negotiate the format, call VIDIOC_REQBUFS to tell the driver how many buffers you want, allocate/map those buffers, queue every one of them with VIDIOC_QBUF, then call VIDIOC_STREAMON to start the hardware. We cover REQBUFS, memory models, and the full queue/dequeue loop in detail in upcoming lectures — this lecture focuses on the piece that has to happen first: format negotiation.
struct v4l2_format on a Modern Kernel
Every ioctl that touches buffer geometry — cropping, requesting buffers, queueing, dequeueing, or querying format — needs to know which kind of stream it is operating on. That is exactly what struct v4l2_format exists for. Its shape has not changed structurally in years, though the set of valid type values has grown as V4L2 gained support for more capture styles:
struct v4l2_format {
__u32 type;
union {
struct v4l2_pix_format pix; /* V4L2_BUF_TYPE_VIDEO_CAPTURE */
struct v4l2_pix_format_mplane pix_mp; /* V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE */
struct v4l2_window win; /* V4L2_BUF_TYPE_VIDEO_OVERLAY */
struct v4l2_vbi_format vbi; /* V4L2_BUF_TYPE_VBI_CAPTURE */
struct v4l2_sliced_vbi_format sliced; /* V4L2_BUF_TYPE_SLICED_VBI_CAPTURE */
struct v4l2_sdr_format sdr; /* V4L2_BUF_TYPE_SDR_CAPTURE */
struct v4l2_meta_format meta; /* V4L2_BUF_TYPE_META_CAPTURE */
__u8 raw_data[200];
} fmt;
};
The type field is set by the application and tells the kernel which member of the union is valid. Almost every modern sensor and ISP driver you will run into today prefers the multiplanar types — V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE using pix_mp — because it lets a single buffer format describe formats where luma and chroma live in separate memory planes, such as NV12M or YUV420M. The classic single-plane V4L2_BUF_TYPE_VIDEO_CAPTURE with pix is still fully supported and remains the simplest starting point, which is what we use in this lecture’s demo. The complete, current list of buffer types lives in the enum v4l2_buf_type definition inside the kernel’s videodev2.h uAPI header — it is worth reading once so you recognize every type when you see it in driver source.
Important note: A modern kernel also exposes V4L2_BUF_TYPE_META_CAPTURE and V4L2_BUF_TYPE_META_OUTPUT for sensor metadata streams (exposure, gain, embedded statistics) that ship alongside the pixel stream on ISP-based platforms — these did not exist when V4L2 buffer negotiation was first documented in older texts, so don’t be surprised to see them in current driver code.
Querying the Current Format with VIDIOC_G_FMT
Applications almost always query the driver’s current format before changing anything, rather than guessing blindly. The rule is simple: pass a freshly zeroed struct v4l2_format with only type filled in, and the kernel fills in the rest.
struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ioctl(fd, VIDIOC_G_FMT, &fmt) == -1) {
perror("VIDIOC_G_FMT");
close(fd);
exit(EXIT_FAILURE);
}
printf("Current format: %ux%u, fourcc=%c%c%c%c\n",
fmt.fmt.pix.width, fmt.fmt.pix.height,
fmt.fmt.pix.pixelformat & 0xFF,
(fmt.fmt.pix.pixelformat >> 8) & 0xFF,
(fmt.fmt.pix.pixelformat >> 16) & 0xFF,
(fmt.fmt.pix.pixelformat >> 24) & 0xFF);
Zeroing the structure matters more than it looks. Any garbage left in the union from the stack can be misread by the driver on some code paths, and it makes debugging painful when fields you never touched appear to have random values.
Common V4L2 Pixel Formats
Once you have the current format, you typically only change the fields you care about and send the whole structure back. V4L2 fourcc pixel formats you will run into constantly on real hardware include:
| Fourcc | Description | Typical use |
|---|---|---|
V4L2_PIX_FMT_YUYV | YUV 4:2:2, packed/interleaved | Most USB UVC webcams |
V4L2_PIX_FMT_UYVY | YUV 4:2:2, byte-swapped packed | Analog/BT.656 style capture bridges |
V4L2_PIX_FMT_NV12 | YUV 4:2:0, semi-planar | Hardware video encoders/ISPs |
V4L2_PIX_FMT_NV16 | YUV 4:2:2, semi-planar | ISP intermediate formats |
V4L2_PIX_FMT_RGB24 | RGB 8:8:8, packed | Simple display pipelines, testing |
V4L2_PIX_FMT_SRGGB10 | 10-bit raw Bayer | Raw sensor output before ISP demosaic |
Setting a New Format with VIDIOC_S_FMT
To change the format, fill in the fields you need and issue VIDIOC_S_FMT. Because some kernels still route this through a compatibility wrapper on certain drivers, using the xioctl()-style EINTR-safe wrapper from our previous lecture is good practice here too.
#define CAP_WIDTH 1280
#define CAP_HEIGHT 720
#define CAP_PIXFMT V4L2_PIX_FMT_YUYV
struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
fmt.fmt.pix.width = CAP_WIDTH;
fmt.fmt.pix.height = CAP_HEIGHT;
fmt.fmt.pix.pixelformat = CAP_PIXFMT;
fmt.fmt.pix.field = V4L2_FIELD_NONE;
if (ep_xioctl(fd, VIDIOC_S_FMT, &fmt) == -1) {
perror("VIDIOC_S_FMT");
close(fd);
exit(EXIT_FAILURE);
}
Notice we did not fill in bytesperline, sizeimage, or colorspace — for VIDIOC_S_FMT on most modern capture drivers, leaving these at zero tells the driver to calculate them for you based on width, height, and pixel format. This is safer than computing them by hand, since only the driver knows the true stride requirements of the underlying hardware.
Important note: A successful return from VIDIOC_S_FMT does not mean your exact request was honored. A device may not support every width/height/pixel-format combination. When that happens, the driver silently substitutes the closest values it does support and writes them back into the same fmt structure. Always re-check the structure after the call.
if (fmt.fmt.pix.pixelformat != CAP_PIXFMT)
fprintf(stderr, "Driver rejected requested pixel format, using its own choice instead\n");
if (fmt.fmt.pix.width != CAP_WIDTH || fmt.fmt.pix.height != CAP_HEIGHT)
fprintf(stderr, "Driver adjusted resolution to %ux%u\n",
fmt.fmt.pix.width, fmt.fmt.pix.height);
Controlling Frame Rate with VIDIOC_G_PARM / VIDIOC_S_PARM
Format negotiation covers geometry and pixel layout, but not timing. Frame rate is a separate negotiation using struct v4l2_streamparm:
struct v4l2_streamparm parm;
memset(&parm, 0, sizeof(parm));
parm.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ioctl(fd, VIDIOC_G_PARM, &parm) == -1) {
perror("VIDIOC_G_PARM");
} else if (parm.parm.capture.capability & V4L2_CAP_TIMEPERFRAME) {
parm.parm.capture.timeperframe.numerator = 1;
parm.parm.capture.timeperframe.denominator = 30; /* request 30 fps */
if (ioctl(fd, VIDIOC_S_PARM, &parm) == -1)
perror("VIDIOC_S_PARM");
}
Check V4L2_CAP_TIMEPERFRAME before attempting VIDIOC_S_PARM — many simple UVC devices don’t support arbitrary frame-rate control at all, and calling it blindly just wastes a round trip. You can enumerate exactly which frame intervals a device supports for a given resolution and pixel format with VIDIOC_ENUM_FRAMEINTERVALS, which we will use directly in the v4l2-ctl lecture later in this chapter.
Demo: ep_v4l2fmt — Original Format Negotiation Tool
Here is a complete, original demo application, ep_v4l2fmt, that opens a capture device, prints its current format, requests 1280×720 YUYV, and reports whether the driver honored the request.
/* ep_v4l2fmt.c — EmbeddedPathashala original demo, not derived from any book source */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <errno.h>
#include <fcntl.h>
#include <unistd.h>
#include <sys/ioctl.h>
#include <linux/videodev2.h>
static int ep_xioctl(int fd, unsigned long req, void *arg)
{
int r;
do {
r = ioctl(fd, req, arg);
} while (r == -1 && errno == EINTR);
return r;
}
int main(int argc, char *argv[])
{
const char *dev = (argc > 1) ? argv[1] : "/dev/video0";
int fd = open(dev, O_RDWR);
if (fd == -1) {
perror("open");
return 1;
}
struct v4l2_format fmt;
memset(&fmt, 0, sizeof(fmt));
fmt.type = V4L2_BUF_TYPE_VIDEO_CAPTURE;
if (ep_xioctl(fd, VIDIOC_G_FMT, &fmt) == -1) {
perror("VIDIOC_G_FMT");
close(fd);
return 1;
}
printf("[%s] Current format: %ux%u fourcc=%c%c%c%c\n", dev,
fmt.fmt.pix.width, fmt.fmt.pix.height,
fmt.fmt.pix.pixelformat & 0xFF,
(fmt.fmt.pix.pixelformat >> 8) & 0xFF,
(fmt.fmt.pix.pixelformat >> 16) & 0xFF,
(fmt.fmt.pix.pixelformat >> 24) & 0xFF);
fmt.fmt.pix.width = 1280;
fmt.fmt.pix.height = 720;
fmt.fmt.pix.pixelformat = V4L2_PIX_FMT_YUYV;
fmt.fmt.pix.field = V4L2_FIELD_NONE;
if (ep_xioctl(fd, VIDIOC_S_FMT, &fmt) == -1) {
perror("VIDIOC_S_FMT");
close(fd);
return 1;
}
printf("[%s] Granted format: %ux%u fourcc=%c%c%c%c\n", dev,
fmt.fmt.pix.width, fmt.fmt.pix.height,
fmt.fmt.pix.pixelformat & 0xFF,
(fmt.fmt.pix.pixelformat >> 8) & 0xFF,
(fmt.fmt.pix.pixelformat >> 16) & 0xFF,
(fmt.fmt.pix.pixelformat >> 24) & 0xFF);
if (fmt.fmt.pix.pixelformat != V4L2_PIX_FMT_YUYV)
fprintf(stderr, "Note: driver substituted a different pixel format\n");
close(fd);
return 0;
}
Build and run it against a real capture device:
$ gcc -Wall -o ep_v4l2fmt ep_v4l2fmt.c
$ ./ep_v4l2fmt /dev/video0
[/dev/video0] Current format: 640x480 fourcc=YUYV
[/dev/video0] Granted format: 1280x720 fourcc=YUYV
If your webcam does not support 1280×720 at YUYV, you will see the granted resolution snap to the closest supported mode instead — that is expected driver behavior, not a bug in the demo. You can cross-check the result independently:
$ v4l2-ctl -d /dev/video0 --get-fmt-video
Format Video Capture:
Width/Height : 1280/720
Pixel Format : 'YUYV' (YUYV 4:2:2)
Field : None
Common Mistakes and Troubleshooting
- Not zeroing the structure: passing a stack-garbage
struct v4l2_formatintoVIDIOC_S_FMTcan leave unrelated union fields set to unpredictable values. - Assuming the request was honored: always re-read
width,height, andpixelformatfrom the structure afterVIDIOC_S_FMTreturns. - Calling S_FMT after STREAMON: format must be negotiated before
VIDIOC_REQBUFS/VIDIOC_STREAMON— most drivers returnEBUSYif you try to change format mid-stream. - Ignoring V4L2_CAP_TIMEPERFRAME: calling
VIDIOC_S_PARMon a device that doesn’t support frame-rate control wastes a call and can be mistaken for a driver bug. - Forgetting multiplanar devices: on an MPLANE-only sensor driver,
fmt.fmt.pixis the wrong union member — you must usefmt.fmt.pix_mpand the typeV4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE, or every ioctl will fail withEINVAL.
Best Practices
- Always query with
VIDIOC_G_FMTbefore setting, so you know the driver’s defaults and only change what you actually need. - Validate the granted format before proceeding to
VIDIOC_REQBUFS— don’t assume your requested resolution stuck. - Prefer letting the driver compute
bytesperline/sizeimagerather than hand-calculating them, unless you have a specific hardware alignment requirement. - Check device capabilities (from the earlier
VIDIOC_QUERYCAPlecture) forV4L2_CAP_VIDEO_CAPTURE_MPLANEbefore assuming single-planepixis correct. - Wrap every ioctl call in an EINTR-safe retry helper — signal interruption is common in real capture applications.
Interview Questions
Why does V4L2 use two separate buffer queues instead of one?
Separating the driver’s incoming (empty) queue from the application’s outgoing (filled) queue lets the kernel and hardware fill buffers asynchronously in strict FIFO order while the application processes previously filled buffers at its own pace, without either side blocking the other unnecessarily.
What happens if VIDIOC_S_FMT is called with an unsupported resolution?
The ioctl still returns success. The driver adjusts the requested width, height, or pixel format to the closest combination it actually supports and writes those adjusted values back into the same struct v4l2_format passed by the caller — it is the application’s responsibility to check the returned values.
How do you know whether to use struct v4l2_pix_format or v4l2_pix_format_mplane?
Check the device’s reported capabilities from VIDIOC_QUERYCAP. If V4L2_CAP_VIDEO_CAPTURE_MPLANE is set, the device expects the multiplanar pix_mp member with buffer type V4L2_BUF_TYPE_VIDEO_CAPTURE_MPLANE; using the single-plane pix member against such a device returns EINVAL.
Summary and Key Takeaways
V4L2 buffer format management rests on two ideas that every capture application must respect: buffers move through a strict two-queue cycle between driver and application, and format negotiation via VIDIOC_G_FMT/VIDIOC_S_FMT is a request, not a guarantee — the driver always has the final say on what it can actually deliver. Getting this right before touching VIDIOC_REQBUFS or streaming ioctls saves hours of confusing debugging later. In the next lecture in this free Linux kernel development course, we extend this into VIDIOC_G_PARM/VIDIOC_S_PARM in more depth and start allocating real buffers with VIDIOC_REQBUFS.
Frequently Asked Questions
What is the difference between VIDIOC_G_FMT and VIDIOC_S_FMT?
VIDIOC_G_FMT queries the device’s current format without changing anything. VIDIOC_S_FMT requests a new format; the driver may adjust the requested values to the nearest supported configuration.
Do I need to call VIDIOC_G_FMT before VIDIOC_S_FMT?
It is not strictly required, but it is good practice — it lets you start from the driver’s defaults and change only the fields you care about instead of guessing every field from scratch.
Why does the driver change my requested width and height?
Hardware capture pipelines only support a fixed set of resolutions and formats. When your request doesn’t exactly match one, the driver rounds to the closest supported combination and reports the actual values back to you.
What is V4L2_FIELD_NONE and why set it?
It tells the driver the stream is fully progressive with no interlacing, which is the correct setting for essentially all modern USB and MIPI CSI-2 camera sensors.
Can I change the format while streaming is active?
No. Format must be set before VIDIOC_REQBUFS and VIDIOC_STREAMON. Most drivers return EBUSY if you attempt VIDIOC_S_FMT once streaming has started.
What does V4L2_CAP_TIMEPERFRAME mean?
It is a capability flag reported via VIDIOC_G_PARM that tells you whether the driver supports changing the capture frame rate at all. If it’s not set, VIDIOC_S_PARM has no effect.
What is the difference between V4L2_PIX_FMT_NV12 and V4L2_PIX_FMT_YUYV?
YUYV is a packed 4:2:2 format with luma and chroma interleaved in one plane. NV12 is a semi-planar 4:2:0 format with a separate luma plane and an interleaved chroma plane — common on hardware ISPs and video encoders.
Is this course free?
Yes. This lecture is part of Embedded Pathashala’s free Linux kernel development course, free Linux device drivers course, and free embedded Linux course material.
Continue Your Free Linux Kernel Development Course
Master the complete V4L2 user space API step by step, with original demos and real kernel behavior — no vendor code, no fluff.
Next Lecture Browse Full Course Index