<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Rumbling Bytes]]></title><description><![CDATA[My small space in the internet to post my not-so-random ramblings about bytes and systems!]]></description><link>https://rumbling-bytes.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/69ce7c830ff860b6dee665b0/05409b95-a7eb-44c7-a4c8-d8f505974654.png</url><title>Rumbling Bytes</title><link>https://rumbling-bytes.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Tue, 29 Sep 2026 07:16:05 GMT</lastBuildDate><atom:link href="https://rumbling-bytes.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[How Programming Languages Turn File Bytes into Readable Data
]]></title><description><![CDATA[In this article we'll discuss how to connect the concept of files as a stream of data to how C and Java read files and streams.
We'll focus on C and Java as they show this concept very clearly.


From]]></description><link>https://rumbling-bytes.hashnode.dev/how-programming-languages-turn-file-bytes-into-readable-data</link><guid isPermaLink="true">https://rumbling-bytes.hashnode.dev/how-programming-languages-turn-file-bytes-into-readable-data</guid><dc:creator><![CDATA[Gatharĩki Ngigĩ]]></dc:creator><pubDate>Mon, 13 Apr 2026 19:17:09 GMT</pubDate><content:encoded><![CDATA[<p>In this article we'll discuss how to connect the <a href="https://rumbling-bytes.hashnode.dev/the-notion-of-files-as-unstructured-streams-of-bytes">concept of files as a stream of data</a> to how C and Java read files and streams.</p>
<p>We'll focus on C and Java as they show this concept very clearly.</p>
<img src="https://cdn.hashnode.com/uploads/covers/69ce7c830ff860b6dee665b0/84dec989-5034-4e74-9230-3feb419b403f.png" alt="" style="display:block;margin:0 auto" />

<h3>From Files to Streams</h3>
<p>Because files are just bytes, the operating system expose them to programs as <strong>streams</strong>.</p>
<blockquote>
<p>A <strong>stream</strong> = data that flows sequentially, byte by byte.</p>
</blockquote>
<p>You can think of it like drinking water through a straw, you don't grab the whole drink at once instead you sip it continuously. Similarly, programs don't load files magically but open a stream and sequentially read bytes.</p>
<h3>Low-level view in C</h3>
<p>C is very close to the OS, so you can see the concept directly.</p>
<p>Opening a file:</p>
<p><code>FILE *f = fopen("data.txt", "rb");</code></p>
<p>Mode <em><strong>rb</strong></em> means:</p>
<ul>
<li><p>r -&gt; read</p>
</li>
<li><p>b -&gt; binary (treat as raw bytes)</p>
</li>
</ul>
<p>Reading bytes:</p>
<p><code>unsigned char buffer[5];</code></p>
<p><code>fread(buffer, 1, 5, f);</code></p>
<p>What we get:</p>
<p><code>buffer = [72, 101, 108, 108, 111]</code></p>
<p>C has no idea that this is "<em>hello</em>". Thus, if we want text, then we must interpret it:</p>
<p><code>printf("%s", buffer);</code></p>
<p>The program decides that bytes represent characters. This is the purest example of <em>unstructured bytes</em>.</p>
<h3>Text is just a convenience layer</h3>
<p>C provides helpers like:</p>
<p><code>fgets(line, 100, f);</code></p>
<p>This feels like "reading a line of text" but internally it is still reading bytes and stopping when it sees byte value <em><strong>10</strong></em> (newline). So the "lines" don't exist in files - programs invent them.</p>
<h3>Java - adds abstraction layers</h3>
<p>Java hides raw bytes behind <strong>InputStreams</strong> and <strong>Readers</strong>. This is where the concept becomes very important.</p>
<blockquote>
<p>FileInputStream = raw bytes</p>
</blockquote>
<p><code>FileInputStream fis = new FileInputStream("data.txt");</code></p>
<p><code>int b = fis.read();</code></p>
<p><strong>b</strong> is an <strong>int 0-255</strong></p>
<p>Again, Java is reading <strong>one byte</strong>.</p>
<p>Example output:</p>
<blockquote>
<p>72</p>
<p>101</p>
<p>108</p>
<p>108</p>
<p>111</p>
</blockquote>
<p>Still just numbers.</p>
<h3>Converting bytes to text</h3>
<p>Humans don't want numbers, we want characters. So Java adds another layer:</p>
<p><code>InputStreamReader reader = new InputStreamReader(fis, "UTF-8");</code></p>
<p>Now Java knows:</p>
<blockquote>
<p>interpret bytes as UTF-8 characters.</p>
</blockquote>
<p>Then we wrap again:</p>
<p><code>BufferedReader br = new BufferedReader(reader);</code></p>
<p><code>String line = br.readLine();</code></p>
<p>Now we can read:</p>
<blockquote>
<p>Hello</p>
</blockquote>
<p>But notice the stack:</p>
<blockquote>
<p>File -&gt; bytes -&gt; characters -&gt; lines</p>
</blockquote>
<p>Multiple layers added by software.</p>
<h3>Why Java has many stream classes</h3>
<p>Because files are just bytes, Java lets you build pipelines. A common chain would be something such as:</p>
<blockquote>
<p>FileInputStream -&gt; raw bytes from file</p>
<p>InputStreamReader -&gt;bytes to characters</p>
<p>BufferedReader -&gt; characters to lines</p>
</blockquote>
<p>This layered design exists because <em>files are unstructured</em>.</p>
<h3>Binary file example (image)</h3>
<p>Reading an image in Java:</p>
<p><code>FileInputStream fis = new FileInputStream("photo.jpg");</code></p>
<p><code>byte[] data = fis.readAllBytes();</code></p>
<p>Java does NOT know it's an image. But this does:</p>
<p><code>BufferedImage img = ImageIO.read(new File("photo.jpg"));</code></p>
<p>Why? Because <strong>ImageIO</strong> understands JPEG format.</p>
<p>Again: OS gives bytes, library interprets bytes.</p>
<h3>Streams are everywhere (not just files)</h3>
<p>Once you understand this, a huge realization happens: these are all byte streams:</p>
<blockquote>
<p>Source Why</p>
<p>Files bytes from disk</p>
<p>Network bytes from internet</p>
<p>Keyboard bytes from user input</p>
<p>Camera bytes from hardware</p>
</blockquote>
<p>This is why Java uses streams for file I/O, sockets, HTTP and pipes. Everything is just <em>flowing bytes</em>.</p>
<p>Deep insight</p>
<p>Hence the famous Unix idea explains everything:</p>
<blockquote>
<p>Everything is a file</p>
</blockquote>
<p>Examples in Linux:</p>
<blockquote>
<p>/dev/keyboard</p>
<p>/dev/mouse</p>
<p>/dev/random</p>
<p>/dev/sda</p>
</blockquote>
<p>Hardware devices behave like files.</p>
<h3>Big takeaway</h3>
<p>When your program reads a file:</p>
<ol>
<li><p>OS gives raw bytes</p>
</li>
<li><p>Language adds abstractions</p>
</li>
<li><p>Libraries interpret structure</p>
</li>
<li><p>Your program gets meaningful data</p>
</li>
</ol>
<p>Flow:</p>
<blockquote>
<p>Disk -&gt; Bytes -&gt; Encoding -&gt; Structure -&gt; Meaning</p>
</blockquote>
<h3>Why this matters in real projects</h3>
<p>You now understand:</p>
<ul>
<li><p>encoding bugs</p>
</li>
<li><p>corrupted files</p>
</li>
<li><p>file parsing</p>
</li>
<li><p>network protocols</p>
</li>
<li><p>serialization</p>
</li>
<li><p>databases</p>
</li>
<li><p>APIs</p>
</li>
</ul>
<p>They all start with <strong>streams of bytes</strong>.</p>
<p>In the next article we will advance this concept and connect it to serialization (<strong>JSON</strong>, <strong>XML</strong>, <strong>protobuf</strong>) which is the next layer above byte streams.</p>
]]></content:encoded></item><item><title><![CDATA[The Notion of Files as Unstructured Streams of Bytes]]></title><description><![CDATA[The idea of files as an unstructured bytes streams was drawn by Ken Thompson when conceptualizing a new operating system from MULTICS. In this article, and perhaps the next, I will try to unpack this ]]></description><link>https://rumbling-bytes.hashnode.dev/the-notion-of-files-as-unstructured-streams-of-bytes</link><guid isPermaLink="true">https://rumbling-bytes.hashnode.dev/the-notion-of-files-as-unstructured-streams-of-bytes</guid><dc:creator><![CDATA[Gatharĩki Ngigĩ]]></dc:creator><pubDate>Fri, 03 Apr 2026 10:55:59 GMT</pubDate><content:encoded><![CDATA[<p>The idea of files as an unstructured bytes streams was drawn by Ken Thompson when conceptualizing a new operating system from MULTICS. In this article, and perhaps the next, I will try to unpack this idea of how the operating system treats files at the lowest level and how that relates to programming.</p>
<blockquote>
<p>TAKEAWAY: The meaning of bytes is decided by the program that's reading the file and not the operating system.</p>
</blockquote>
<img src="https://cdn.hashnode.com/uploads/covers/69ce7c830ff860b6dee665b0/1d791333-a38d-4700-9563-f2b9ba4176a9.png" alt="" style="display:block;margin:0 auto" />

<h3>The beginning: What is a byte</h3>
<p>A byte = 8 bits = a number between <strong>0 to 255</strong></p>
<p>Example byte sequence:</p>
<blockquote>
<p>72 101 108 108 111</p>
</blockquote>
<p>Only when you interpret it in <a href="https://en.wikipedia.org/wiki/ASCII"><em>ASCII</em></a> format does the byte sequence become:</p>
<blockquote>
<p>Hello</p>
</blockquote>
<p>But the OS does NOT know this is text. It only sees the numbers.</p>
<h3>Why "unstructured"?</h3>
<p>Well, because the OS does not impose any format or structure such as lines, words, images, tables, records, etc,.</p>
<p>Those structures only exist in file formats, which are basically software created conventions and so, the same bytes can mean completely different things depending on the program reading them.</p>
<p>For instance, consider these bytes:</p>
<blockquote>
<p>50 51 10</p>
</blockquote>
<p>Possible interpretations:</p>
<p><strong>Program reading it Meaning</strong></p>
<p>Text editor Characters "23\n***"***</p>
<p>Image viewer Corrupted image</p>
<p>Music player Noise</p>
<p>Compiler Source code fragment</p>
<p>As you can see the OS doesn't care - it just stores bytes.</p>
<h3>Text vs Binary files</h3>
<p>Now, people will often say text files and binary files. From the OS perspective there is no difference, both are just bytes. The difference only exist in how software interprets them.</p>
<p>For example:</p>
<p><strong>File Reality</strong></p>
<p><em>.txt</em> bytes interpreted as characters</p>
<p><em>.jpg</em> bytes interpreted as image encoding</p>
<p><em>.mp3</em> bytes interpreted as sound encoding</p>
<p><em>.exe</em> bytes interpreted as machine instructions</p>
<h3>Example: A picture file</h3>
<p>A JPEG image is not "a picture" on disk. It's just bytes like:</p>
<blockquote>
<p>FF D8 FF E0 00 10 4A 46 49 46 ...</p>
</blockquote>
<p>When an image viewer reads these bytes, it:</p>
<ol>
<li><p>Recognizes JPEG format</p>
</li>
<li><p>Decodes bytes</p>
</li>
<li><p>Renders pixels</p>
</li>
</ol>
<p>And so without the viewer, it's just meaningless numbers.</p>
<h3>Why OS uses this design</h3>
<p>This design is powerful and flexible.</p>
<ol>
<li>Simplicity</li>
</ol>
<p>The OS just needs to provide basic operations:</p>
<ul>
<li><p>Open file</p>
</li>
<li><p>Read bytes</p>
</li>
<li><p>Write bytes</p>
</li>
<li><p>Close file</p>
</li>
</ul>
<p>Thus removes the need to understand or support thousands of formats.</p>
<ol>
<li>Portability</li>
</ol>
<p>Programs can just define their own formats and they are guaranteed to work across systems.</p>
<p>In essence, any program can store any kind of data such as text, video, databases, AI models or compressed archives, all using the same interface.</p>
<h3>How programs add structure</h3>
<p>Programs impose structure using <em>file formats.</em> Here is are examples of structures added on top of raw bytes:</p>
<p>Example 1: Text file structure</p>
<blockquote>
<p>Hello\nWorld</p>
</blockquote>
<p>Structure added:</p>
<ul>
<li><p>Characters</p>
</li>
<li><p>Lines</p>
</li>
<li><p>Encoding (UTF-8)</p>
</li>
</ul>
<p>Example 2: CSV structure</p>
<blockquote>
<p>Name,Age</p>
<p>Alice,25</p>
<p>Bob,30</p>
</blockquote>
<p>Structure added:</p>
<ul>
<li><p>Rows</p>
</li>
<li><p>Columns</p>
</li>
<li><p>Delimiter rules</p>
</li>
</ul>
<p>Example 3: Database file</p>
<p>Structure added:</p>
<ul>
<li><p>Pages</p>
</li>
<li><p>Indexes</p>
</li>
<li><p>Schemas</p>
</li>
<li><p>Records</p>
</li>
</ul>
<p>All still just bytes underneath.</p>
<h3>Why this matter in programming</h3>
<p>This simple yet noble concept explains many real behaviors such as:</p>
<ul>
<li>Why file extensions don't matter</li>
</ul>
<p>Renaming: <em><strong>photo.jpg -&gt; photo.txt</strong></em></p>
<p>doesn't change the file - only how programs try to interpret it.</p>
<ul>
<li>Why corrupted files exist</li>
</ul>
<p>If some bytes change, then the program cannot correctly interpret structure and thus the file appears broken.</p>
<ul>
<li>Why encoding issues happen</li>
</ul>
<p>Consider these bytes: <em><strong>C3 A9</strong></em></p>
<p>in UTF-8 = "é"</p>
<p>in ASCII = "Ã©"</p>
<p>Same bytes, different interpretation.</p>
<h3>Summary</h3>
<p>In short, we can conclude that files are:</p>
<ul>
<li><p>Just sequence of bytes</p>
</li>
<li><p>With no built-in structure</p>
</li>
<li><p>The OS is format-agnostic</p>
</li>
<li><p>Structure is defined by software that reads the file</p>
</li>
</ul>
<p>In the follow-up article, I'll try practically connect this very concept to how Java and C read files and streams.</p>
<p><strong>References</strong><br />"... and the notion of files as unstructured streams of bytes" (Kerrisk, <em>The Linux Programming Interface</em>, p. 2).</p>
]]></content:encoded></item></channel></rss>