The Ridiculous Engineering Of YouTube Videos
Source: The Ridiculous Engineering Of YouTube Videos, Enrico Tartarotti, 14:57, uploaded 2026-05-07, playlist index 48.
YouTube presents a page with a thumbnail, a title, and a play button. Tartarotti takes that small surface apart and finds three systems underneath it. The first system engineers the interface around playing time. The second engineers the video around attention and choice. The third turns a single upload into a large set of files that can reach a particular device, connection, and moment. The apparent simplicity belongs to the viewer. The work sits underneath.
Friction around choosing and continuing
The opening example is the playable thumbnail. On desktop, hovering over a video starts a preview. On mobile, waiting a few seconds does the same. The preview includes the title, the thumbnail, several seconds of footage, a seek bar, subtitles, and an audio control. Tartarotti says he sometimes watches minutes or whole videos in this small window. YouTube treats a preview that lasts a few seconds like a regular view and saves it to watch history.
He dates the current version of this interface to 2021 and gives its reason as a product problem: choosing a video creates one of the highest points of friction in the experience. The preview lets a person inspect a video without opening its page or making a full commitment. The result is a smaller gap between browsing and playing.
The same idea appears in the mobile player. Closing a video takes two presses. The first press sends it into picture-in-picture mode whilst the person continues looking for something else. A second press on the X removes it. Swiping down keeps playback in picture-in-picture mode, and swiping up opens the video full-screen. Both gestures preserve playback. On television, recommendations can fill most of the screen whilst the original video continues underneath them. Tartarotti calls this friction engineering.
His more precise term is non-playing time. YouTube wants to reduce the period in which a person is present in the app without watching anything. He contrasts this with the twenty minutes a Netflix user can spend searching for something to watch. That time counts as app use, yet it can leave the user with the sense that the service has nothing worth watching. Tartarotti says YouTube therefore optimises for playing time rather than time spent inside the app. Playable thumbnails and continuing playback grow the first measure whilst shrinking the second.
This creates a design problem that did not exist in the first YouTube. The early service had about forty videos and needed a way to fill a player. Short-form platforms later trained people to receive an automatically playing stream instead of choosing each item. YouTube now has to preserve the value of a large catalogue and deliberate choice whilst adapting to a habit formed by TikTok, Instagram, and Shorts.
Tartarotti shows a 2024 clip posted by a user on X in which a test version of YouTube lets people scroll through full-screen long-form videos. The gesture resembles an endless short-form feed, although the videos remain long. He says the test never shipped and uses it as evidence of the scale of YouTube’s experimentation. The platform can run many A/B tests at once and observe how people respond before it commits to a redesign.
PewDiePie’s attempt to restore an older YouTube with browser extensions supplies the counterexample. He removes Shorts, the algorithmic home page, and newer interface features. Tartarotti says he wants that version too, and then applies the product manager’s rule he learned from building technology: observe behaviour rather than taking stated preference as a reliable guide. The five-second autoplay countdown, reduced from ten seconds in 2014, survives custom scripts and user complaints because the behaviour data tells YouTube that it works. Tartarotti finds the newer quality picker, which replaces exact resolutions with labels such as higher quality and lower quality, infuriating. He can still see the reason for hiding a technical choice from most viewers.
The platform keeps testing the boundary between useful choice and an endless feed. Its interface has to make a video easy to start without making YouTube feel like a stream of interchangeable fragments.
The thumbnail as a piece of content
The second layer turns the camera back on the video itself. Tartarotti points out that the hook, the reveal, and the order of his examples all came from planning. He shows the picture-in-picture control before explaining what makes it strange. The apparent ease of the presentation is produced through the same kind of decisions that shape the platform around it.
He argues that YouTube’s strength comes from combining a large choice of high-quality, in-depth material with a system that lets a person choose what to watch. That choice also creates a severe competition for the first click. One of his videos received 3,776 views across its first 55 days and then reached 3.5 million views after he changed the thumbnail. The video stayed the same. For that upload he tried twenty thumbnails, and he shows sheets of variations for other videos as well.
The thumbnail rules he gives are practical. Dark mode makes lighter backgrounds stand out. Red attracts attention when it contrasts with surrounding colours. Realism matters less than recognition within roughly 0.1 seconds. The drawings in this video’s thumbnail exaggerate the subject and suggest that something about YouTube remains hidden. Their role is to make the subject legible and create a reason to click. Tartarotti cites a full test by MrBeast in which a closed mouth performed slightly better than an open mouth. MrBeast now keeps his mouth closed in thumbnails, according to the video. Creators change shirt colours, positions, and other small details across many versions to improve the chance of a click.
The old ideal of an organic YouTube video receives a direct test. Tartarotti shows what the title and thumbnail would look like if they described the video plainly and kept the older style. He then asks which version a viewer would choose. The point is uncomfortable because the engineered version often communicates the subject faster. Once every video competes inside an attention system, the visible work absorbs the logic of the platform around it.
Content engineering continues after the click. YouTube Studio shows a retention curve that maps where viewers leave, stay, or respond to particular phrases. Tartarotti has added extensions that expose more measures, and he compares the resulting dashboard with serious analytics software. He then asks what happens when the data becomes too heavy for the creator.
His example is the comments page in YouTube Studio. In an earlier video, he calculated that 21.4 per cent of comments on his channel were positive, 17.1 per cent neutral, and 61.5 per cent negative. He connects the distribution to negativity bias: a bad restaurant experience is more likely to produce a one-star review than a good one is to produce a five-star review. He says YouTube Studio began showing him mostly positive comments in its overview, whilst the expanded list still contained negative ones. He infers that YouTube filters the first view to shield creators from unnecessary negativity, then states that he cannot confirm the change because YouTube has not documented it.
That qualification matters to his wider claim. YouTube needs creators to keep making videos, and it needs new creators to have a plausible chance of finding an audience. Tartarotti presents the newer Hype feature as one response. A viewer can give props to a video, and the algorithmic boost is said to be inversely related to the channel’s subscriber count. A small channel receives more help from the same action. The video gives the mechanism as a platform claim; it provides no public experiment or documentation for the exact relationship.
A video as a delivery system
The third layer begins with the click. When the viewer opens this page, the browser sends YouTube a request containing the user’s ID, browser information, cookies, and the video’s ID. Tartarotti’s point is that YouTube does not send back one video file because the upload has already become a collection of files.
He uses a hypothetical 20 GB 4K upload to explain the transformation. YouTube creates versions at 4K, 1080p, 720p, 480p, and 360p, with the 4K stream itself encoded below the quality of the original file. It then cuts each version into segments that last roughly two to ten seconds. One upload becomes hundreds of small files.
The first response is a DASH manifest. Tartarotti compares it to a restaurant menu because it lists the available versions and the links to them. The browser reads the screen resolution, current connection speed, and device, then asks for the first segment of a chosen version. A two-second piece can arrive and start playing whilst the player requests the next piece and checks the connection again. The grey line beside the playhead shows the segments already downloaded and ready to play.
Adaptive quality changes the next request. When the connection slows during a large download, the browser can request the next segment at a lower quality. The video then gives an encoding ladder as another hidden layer of platform economics. A new video with zero views starts in H.264, which Tartarotti describes as the cheapest and lowest-quality encoding. Around 3,000 views, he says, YouTube re-encodes it to VP9. At one million views, it receives AV1, the highest tier in his account. Popularity therefore changes the quality of the file available to the audience. He asks viewers to leave a Hype if they want this video to reach the higher encoding tier.
This is a simplified account. Tartarotti says it leaves out CDNs, codecs, and other pieces of the delivery system. The simplification still gives the right shape: a video that appears to arrive as one continuous object is assembled through repeated requests, quality decisions, and background encoding.
The scale makes the system harder to think about. Tartarotti says YouTube serves 106,000 years of watched video every day. He then returns to the company’s early history. The founders began with roughly forty videos and soon faced the opposite problem: delivery costs grew faster than the young business. He reports annual revenue of “1 billion and says that each view cost about one cent to deliver. The caption’s “$30 a year” is internally implausible and may have dropped a unit such as million, so I leave that figure unresolved rather than silently correcting it.
The video places those figures around Google’s 2006 acquisition of YouTube for $1.65 billion. In 2007, Tartarotti says, YouTube consumed as much bandwidth as the entire internet had used in 2000. By 2014, he says, it accounted for 20 per cent of all internet traffic measured in bits. The historical figures support his larger point about infrastructure, although the video does not identify the underlying reports.
Google supplied the data centres, networks, and servers that a video platform needed in 2006. Tartarotti calls YouTube’s present scale one of the hardest problems in technology and points to a custom chip built for video encoding. Its job is to compress the upload into the many small segments that can play on a bus in a remote mountain in Peru without a spinning wheel. The viewer sees a play button and a continuous image. The system underneath has to keep deciding which small file can arrive next.
Evidence limits
The video’s interface examples and Tartarotti’s own thumbnail and retention examples document a working practice. They do not establish that every YouTube surface has the motive he assigns to it, or that the reported design choices produce the same effect for every viewer. The 2024 long-form scrolling test, the comment filtering behaviour, the Hype boost, the encoding thresholds, the traffic figures, and the infrastructure history appear as reported examples without linked supporting sources in the description.
The captions are automatic and contain at least one material uncertainty in the reported revenue figure. The delivery explanation also skips the CDN and codec details that would be needed for a technical account of YouTube’s current architecture. Device, account, region, and later product changes can alter the observed interface. The durable claim is the relation between layers: a visible YouTube video depends on interface decisions, creator-side optimisation, and a delivery system that remains hidden because it usually works.
Related: notifications as behaviour design, social media and short-form stories, good information is expensive.