[{"content":"The 1975 ACM Turing Award Lecture.\n","permalink":"https://nachooglpz.pages.dev/archives/011-computer-science-as-empirical-inquiry-symbols-and-search/","summary":"Newell and Simon present their view on computer science: the machine as the organism we study.","title":"Computer science as empirical inquiry: symbols and search","topics":["Research","Heuristics","Software"],"type":"Paper"},{"content":"In my previous entry I wrote about how databases that implement a cost-based optimizer use heuristic pruning (External link) to avoid an exhaustive search of the plan space (which in a query optimizer represents the set of possible ways to implement a query).\nSome of the common heuristics that are implemented in databases (External link) can be:\nSelection cascade and pushdown Applying selections as soon as you have the relevant columns, as this will lower the size of inputs to the joins. Another heuristic that developed this one is the assumption that selection is cheap and/or free, while joins are expensive.\nProjection cascade and pushdown Keeping only the columns that are needed to evaluate downstream operators. While we might need all the columns to do joins, it is better to get rid of them as soon as we don\u0026rsquo;t need them. And this way, sort-merge (External link) or hashing (External link) take less space.\nAvoiding Cartesian products Given the a choice, do theta-joins (External link) rather than cross-product (External link). And although this approach might not always be a good solution, it allows the optimizer to keep the plan space small.\nBut these examples are the heuristics that the query planner of a database might use to discard options in the plan space. Heuristics can be found all over computer science disciplines to describe techniques, fundamentals, and theories on how to design and code programs.\nEmpirical Inquiry The etymology (External link) of the word heuristic means \u0026ldquo;serving to discover or find out\u0026rdquo;, and computer science, as Allen Newell and Herbert Simon might argue (External link), is just another empirical discipline. The scientists, in their 1975 Turing Lecture, presented a thesis for computer science as the discipline to study the machine from an empirical perspective. Their argument lies in that, unlike other sciences, some of the forms of observation and experience do not fit the stereotype of the experimental methods.\nIn computer science each new machine that is built, as well as each new program developed for that proposed machine is an experiment. We observe the machine in operation and analyze its behavior by analytical and measurement means. The most interesting part is that both hardware and software are artifacts that have been designed, and therefore they are not black boxes since we can relate their structure to their behavior and draw lessons from experimenting.\nQualitative Structure Newell and Simon also argued that all sciences start by characterizing qualitatively the systems they study. Biology uses cells; geology, plate tectonics; medicine, germs as a theory for disease; and atoms offer a structure to describe the nature of physics. Laws of qualitative structure are not only everywhere in science, but even some of the greatest scientific discoveries are to be found among them.\nSearch Engine Wars The need of actively developing and deriving qualitative structures for the development of programs and systems can still be observed outside of theory. One of the most important systems used since the boom of the internet are information retrieval (IR) systems as the core of search engines. Google started from the development of PageRank (External link), a system that interpreted the importance of a page by how many other pages linked to it. User preference through the passage of time proved Google\u0026rsquo;s algorithms more effective than the ones implemented by their competition, and even the fall of Yahoo is attributed to their turning down of Google\u0026rsquo;s offer to be bought several times.\nThe performance of retrieval systems (External link) is closely related to the use of heuristics. Research in the IR discipline often focuses on developing heuristics on which algorithms and strategies allow to find results closer to what people need or expect, and since the development of these heuristics are what drive competition, companies (more often than not) keep these research extremely protected, as it is the basis of any IR product. To this day Google noticeably still works on IR research (External link) to keep on improving their search engine.\nIt is through the study of design and implementation of systems for machines, as well as developing the heuristics that pave the algorithms that Google (and other IR products) implement, that these companies and products maintain user preference.\n","permalink":"https://nachooglpz.pages.dev/blog/006-heuristics-are-the-secret-sauce/","summary":"In computer science, each machine and software is an experiment.","title":"Heuristics Are the Secret Sauce","topics":["computer-science","research","heuristics"],"type":"essay"},{"content":"Selinger (et al.) paper on the cost-based optimizer.\n","permalink":"https://nachooglpz.pages.dev/archives/010-access-path-selection-in-a-relational-database-management-system/","summary":"Selinger (et al.) paper on the cost-based optimizer.","title":"Access Path Selection in a Relational Database Management System","topics":["Database"],"type":"Paper"},{"content":"In 1972, Edgar Frank \u0026ldquo;Ted\u0026rdquo; Codd published a paper called \u0026ldquo;Relational Completeness of Data Base Sublanguages\u0026rdquo; (External link), which established an equivalence in expressive power between relational calculus (External link) and relational algebra (External link). The theorem that he developed implied that there was a connection between the declarative representation of queries and their operational description, which opened the door to SQL as we know it.\nThe declarative nature of SQL, allowing a user to state what information they want to retrieve and not how to get it, enables a system to optimize how to retrieve and process the information the user is requesting.\nProgram synthesis In computer science, there is a field called program synthesis (External link), which studies how to automatically generate a program given a formal specification; in other words, this field allows a user to create a program without actually coding it. This idea might remind us of the concepts of vibe coding (External link) or agentic engineering (External link), where the user might specify a program\u0026rsquo;s goal (or even structure) in natural language, although these might not exactly count as program synthesis since English is not a formal language. Nonetheless, as a formal language, SQL might be an example of how databases have been implementing program synthesis for a few decades.\nThe query optimizer During the \u0026rsquo;70s and \u0026rsquo;80s, commercial databases that adopted SQL (or an SQL-like API) had to implement a query optimizer (External link) for the system to plan how to effectively (and efficiently) retrieve the data that the user asked for.\nThe first approach was to use rule-based optimization (External link), allowing the system to design a query execution plan based on a set of rules, such as whether the data to be retrieved had an index, but not accounting for any current state of the database or the data layout. A few years after Ted Codd published his papers justifying his theorem and the proposal for the relational data model, the research division at IBM started developing System R as a relational database. One of the innovations of System R was the implementation of a cost-based optimizer (External link), developed by Patricia Selinger et al. (External link).\nSystem R\u0026rsquo;s cost-based optimizer has come to be considered a framework known as the bible of query optimization. This framework uses catalog stats to find the least-cost plan per query block. The idea is that since each block of a query can be converted to relational algebra, and each operator has different implementation options, operators can be applied in different orders.\nThere are three main components of a cost-based query optimizer:\nThe plan space, which analyzes the possible ways to implement a query. The cost estimation, which analyzes the cost of a particular plan. The search strategy, which aims to answer the question: given that we have a plan space, how do we efficiently find a plan with the cheapest cost estimate? Search problems While there are many foundations of AI, search algorithms (External link) allow systems to find solutions by exploring paths and traversing states in a problem space. Databases that implement a cost-based optimizer implement these search algorithms as part of their search strategy, exploring the different plans of the plan space while trying to find the cheapest path to run a query. And because the plan space might grow exponentially depending on the complexity of a query, modern databases and optimizers also implement the use of heuristic pruning (External link) to avoid an exhaustive search of the plan space.\nQuery optimizers are the magic that allows for databases to translate a formal language such as SQL into a complete data description and manipulation program. The long-standing objective of databases has been to retrieve data and apply operators on said data efficiently, and this objective has driven innovations that incorporate research from other areas of computre science, serving as an example of collaboration across fields.\n","permalink":"https://nachooglpz.pages.dev/blog/005-how-databases-did-ai-before-it-was-cool/","summary":"Databases are an example of program synthesis, while a query optimizer might use heuristic search.","title":"How Databases Did AI Before It Was Cool","topics":["databases","search-problems","AI"],"type":"essay"},{"content":"Edgar Frank “Ted” Codd\u0026rsquo;s paper which established the equivalence in expressive power between relational calculus and relational algebra.\n","permalink":"https://nachooglpz.pages.dev/archives/009-relational-completness-of-data-base-sublanguages/","summary":"Ted Codd\u0026rsquo;s paper which established the equivalence in expressive power between relational calculus and relational algebra.","title":"Relational Completeness of Data Base Sublanguages","topics":["Database","Research"],"type":"Paper"},{"content":"A collection of Bjarne Stroustrup\u0026rsquo;s Quotes.\n","permalink":"https://nachooglpz.pages.dev/archives/008-bjarne-stroustrup-quotes/","summary":"A collection of Bjarne Stroustrup\u0026rsquo;s Quotes.","title":"Bjarne Stroustrup Quotes","topics":["Compilers","Data Structures","Open Source","Performance","Programming","Software"],"type":"Site"},{"content":"The process of execution for the code that we create can go down multiple paths. The processing unit that runs the instructions that our code generates usually uses a technique called pipelining (External link), which grants the CPU the ability to fetch upcoming instructions from memory when executing the current one, allowing it to complete one instruction per clock cycle.\nConsidering that the CPU has three basic instruction cycles (External link) (adding an optional write back cycle), pipelining is based on the notion that the CPU can fetch the next set of instructions even if there is a conditional pending, as the CPU will decide which is the branch with the highest probability to be executed.\nWhen the CPU predicts the branch of a conditional statement correctly, it gains the advantage that it has already fetched the instructions and can run them when the conditional evaluates to what the CPU predicted. But if the CPU is wrong, and the branch of a conditional statement evaluates differently from what the CPU predicted, the CPU needs to fetch the instruction set for the other conditional branch from scratch. It is here where cache hits and misses (External link) impact the performance for CPU prediction.\nAvoiding conditionals Neal Ford, Mark Richards, Pramod Sadalage, and Zhamak Dehghani once wrote (External link):\nDon’t try to find the best design in software architecture; instead, strive for the least worst combination of trade-offs.\nThere are cases where we need the CPU to be as efficient as possible and always fetch the next instructions to run, without having to jump or branch across instructions. And this is where Branchless Programming advocates come to make their argument (External link):\nIf you guess wrong too often, you spend a lot of time stalling, rolling back, and restarting.\nThis approach of avoiding conditionals and using clever techniques (External link) to compute results while avoiding CPU cache misses, relies on the assumption that we know the optimal mechanism for our code to run in the CPU. But there are other people that argue that we are usually not the best at optimizing our own code and execution.\nCode optimization For decades, compiler engineers have dedicated their time to optimize the translations that compilers make from source code to assembly and machine code. Compilers often implement vectorization (External link), jump threading (External link), dead code elimination, and other optimization techniques where branchless conversion also takes part. This means that even when a programmer codes conditionals, the compiler might identify a branchless optimization and apply it to the program, allowing the compiler algorithm to decide if the optimization is actually worthwhile (although these are often based on heuristics, for which the programmer still needs to decide to trust).\nDifferent CPU architectures also implement their own methods to allow for optimizing code, and even compilers can optimize (External link) to take advantage of specific instruction provided by these architectures that allow to execute code without disturbing the branch predictor as the control flow would remain linear.\nThe proponents of not implementing branchless programming argue that this programming paradigm is actually a pessimization (External link), meaning that it is a technique that people apply hoping to improve the performance of a program, when in practice it might hurt it.\nSince the compiler often knows more about the architecture on which a program will run, and probably has better heuristics for code optimization, unless the programmer is anything close to Mel (External link), it might be better to let the compiler optimize the program.\nIn a discussion thread about branchless programming (External link), a user posted their experience noting that the only times they have seen an improvement in performance by using branchless programming is in math-dominated code running on Intel architectures.\nThis might come from the notion that branches can be unpredictable in math-dominated code (compressors might be an example that comes to mind), and it might be difficult for a compiler to optimize calculations for unpredictable conditions in these types of programs.\nExperimenting with performance If improving performance of programs is the final goal of branchless programming, Bjarne Stroustrup might have something to say about the discussion. The first recommendation that he might issue goes in the line of allowing compilers to optimize your code for you, as he has famously said:\nPlease remember not to optimize without measurement showing a need to.\nBut once a need to optimize a program has been identified, it might also be important to determine whether branches are the performance bottleneck, following another quote:\nMeasure, if there is anything that you can measure to get feedback.\nAnother interesting article (External link) measured and benchmarked the performance of conditional and branchless implementations of common computational operations (swap, abs, min, max, and power of two). The results showed that most of the operations had the same performance as their branchless implementations when compiled without optimizations. The author had written:\nBranchless techniques significantly outperform their if-based counterparts when compiled without optimizations. This happens because the compiler does not yet apply aggressive optimizations, leaving conditional branches as costly jumps.\nThe only implementation in which the branchless version was still faster than the conditional was the swap operation, and even there, compiler optimizations made the branchless implementation around 50% faster, suggesting that compiler optimizations still affect branchless code, and this can have either positive or negative effects on performance. Therein lies the importance of measurement.\n","permalink":"https://nachooglpz.pages.dev/blog/004-the-case-for-branchless-programming/","summary":"In the debate around branchless programming, the best approach is to benchmark.","title":"The Case for Branchless Programming","topics":["compilers","optimization","software-design","performance","software"],"type":"essay"},{"content":"A folk story on computer science.\n","permalink":"https://nachooglpz.pages.dev/archives/007-the-story-of-mel-a-real-programmer/","summary":"A folk story on computer science.","title":"The Story of Mel, a Real Programmer","topics":["Compilers","Performance","Programming","Software"],"type":"Site"},{"content":"Andy Pavlo\u0026rsquo;s lecture on his paper What Goes Around Comes Around\u0026hellip; And Around\u0026hellip; (External link).\n","permalink":"https://nachooglpz.pages.dev/archives/006-lecture-what-goes-around-comes-around-and-around/","summary":"Andy Pavlo\u0026rsquo;s lecture on his paper \u0026ldquo;What Goes Around Comes Around\u0026hellip; And Around\u0026hellip;\u0026rdquo;.","title":"Lecture of What Goes Around Comes Around... And Around...","topics":["Database"],"type":"Lecture"},{"content":"Snapshot of the post written by David J. DeWitt and Michael Stonebraker on their views of MapReduce.\n","permalink":"https://nachooglpz.pages.dev/archives/003-mapreduce-a-major-step-backwards/","summary":"David J. DeWitt and Michael Stonebraker\u0026rsquo;s views on MapReduce.","title":"MapReduce: A major step backwards","topics":["Cloud","Database"],"type":"Snapshot"},{"content":"This paper provides a summary of 35 years of data model proposals, grouped into 9 different eras.\n","permalink":"https://nachooglpz.pages.dev/archives/004-what-goes-around-comes-around/","summary":"Sonebraker and Hellerstein\u0026rsquo;s summary of 35 years of data models.","title":"What Goes Around Comes Around","topics":["Database"],"type":"Paper"},{"content":"This paper is a continuation of the What Goes Around Comes Around (External link) paper by Stonebraker and Hellerstein. The authors argue that relational model continues to be the dominant model and SQL as been extended to capture the good ideas from other data models.\n","permalink":"https://nachooglpz.pages.dev/archives/005-what-goes-around-comes-around-and-around/","summary":"Stonebraker and Pavlo\u0026rsquo;s analysis of 20 years of data models.","title":"What Goes Around Comes Around... And Around...","topics":["Database"],"type":"Paper"},{"content":"With the boom of the Internet in the early 2000s, the amount of data Google had to process exploded. They had designed a distributed file system (External link) in which large files are partitioned, distributed, and replicated across nodes for fault tolerance.\nTo be able to process all of this data, they also developed MapReduce (External link), a high-level programming model for parallel data processing which, as its name implies, is composed of two phases: Map and Reduce. As a leader in large-scale data processing, Google\u0026rsquo;s file system design and processing paradigm quickly became attractive; and an open-source version of it, known as Hadoop (External link), quickly became popular.\nA restrictive paradigm In 2008, David J. DeWitt and Michael Stonebraker wrote about their reactions (External link) to the MapReduce paradigm, arguing that, although the paradigm may have been a good idea for writing certain types of general-purpose computations, it was a step backwards for the database community.\nThe following are the concerns that they noted.\nDatabase access and schema The scientists argued that schemas are a crucial part of an application program, as they allow systems to keep \u0026ldquo;garbage\u0026rdquo; out of a dataset by providing a way for the runtime to ensure that input records obey said schema.\nJoe Hellerstein has argued (External link) that we need to separate the layers of a system and provide independence between those layers, creating what he called \u0026ldquo;Hellerstein\u0026rsquo;s Inequality\u0026rdquo;:\nData independence is most important when the rate of change of your environment exceeds the rate of change of your applications.\nDeWitt and Stonebraker argue that MapReduce forces the access program to specify an algorithm for data access, instead of stating what information is needed, violating the principle of separating the schema from the application program. They compare this approach to the debate between relational vs. CODASYL, where relational promoted access through a declaration of the desired data, and CODASYL promoted access through the specification of the algorithm to retrieve said desired data.\nThere might be an argument that claims that the datasets that MapReduce targets have no schema; but the scientists refute this by arguing that when extracting a key from the input data set, the map function relies on the existence of at least one data field in each input record.\nBrute force implementation As a second concern, the scientists argued that MapReduce only provides a brute-force processing approach, relying only on providing parallel execution on a grid of shared-nothing computing nodes. What the scientists argue is that this feature has already been implemented in different prototypes for modern database management systems; and that DBMSs also implement hash or B-tree indexes to accelerate access to data which, combined with a query optimizer that decides when to use said index and when to perform a brute-force sequential search, provide a better processing approach.\nSkew factor Another disadvantage of MapReduce is the skew factor. DeWitt and Jim Gray had written (External link) about how skew is an impediment to achieving successful scale-up in parallel query systems. When the distribution of records per key occurs in the map phase, some reduce instances receive more records than others, and the total time of computation depends on the running time of the slowest instance (or the one that receives the most records).\nAlso, when using a pull design (External link), it is inevitable that two reduce nodes will try to access the same map node, reducing the disk transfer rate of that map node.\nAn old paradigm Besides programmers looking at Google and thinking that if Google was using MapReduce, everybody should be using it, the MapReduce community seemed to feel that they had discovered a new paradigm. But the idea of partitioning a large data set into smaller partitions was first proposed in 1983 (External link), with many other implementation proposals following.\nThey also noted that if the ability to write MapReduce functions was what differentiated the paradigm from a parallel SQL implementation, PostgreSQL supported user-defined functions (External link) and user-defined aggregates (External link) since the mid-1980s. Modern database systems have provided such functionality since 1995.\nMissing features and incompatibility DeWitt and Stonebraker mentioned how a modern SQL database has diverse classes of tools available that MapReduce cannot use and for which it has no equivalents of its own. Some of these useful tools are the following.\nReport writers to prepare reports for human visualization. Business intelligence tools to enable querying of data warehouses. Data mining tools to allow users to discover structure in large data sets. Replication tools to allow users to replicate data from one DBMS to another. Design tools to assist the user in building the database. Considering the database access concern, integrity constraints and referential integrity are features to help keep garbage out of the database that were still missing from MapReduce. Views were also missing, so that the schema can be updated.\nThe scientists also noted other features provided by modern DBMSs that are missing in MapReduce. These include a bulk loader to transform input data into a desired format; but one could argue that the objective of the MapReduce paradigm is to be a bulk loader. They also noted the necessity to have indexing (as mentioned previously), updates, and transactions; but one could also argue that if MapReduce is more about the parallel processing of data, these should be implemented elsewhere in the system.\nEven if MapReduce were to be only the parallel processing paradigm of a bigger system, being SQL-incompatible was a big limitation on the use of the mentioned tools.\nDataflow engines The MapReduce paradigm is actually a specific instance of a group of execution systems known as dataflow engines (External link). The base logic for these systems lies in modeling the flow of data through several processing stages: input is partitioned and processed in parallel, then sent to a set of nodes; the output generated by these nodes is then sent through the network to another set of nodes. The map and reduce functions that the MapReduce paradigm uses are also known as operators, where the output of one operator becomes the input of another, and the nodes of a system are characterized by the operators they implement.\nAn example of another dataflow engine is Spark (External link) (which also won the 2022 SIGMOD Systems Award (External link)). This system uses a semi-structured data model, so objects can be anything from key-value pairs to objects of a certain type. These objects are organized into Resilient Distributed Datasets (RDDs) (External link), which are distributed, immutable data sets, together with their lineage (an expression describing how the data was computed).\nWhat is most interesting about this system is that it is based on relational operators (External link) instead of explicit map and reduce functions defined by the user, and therefore it implements its own SQL-like API for users to operate on data.\nComing back to SQL There are two famous papers on database systems that talk about the history of data models (External link); although the one that concerns this entry is the continuation paper (External link), which also discusses MapReduce and dataflow engines overall. The main thesis of these two papers is that data models (and therefore database systems) have followed a cycle (External link) of trying to replace the relational data model and SQL, only to return, with SQL incorporating ideas from the data models that tried to replace it (we could see this behavior in Spark).\nPavlo and Stonebraker had written that at the time of the development of MapReduce, Google had little expertise in DBMS technology, and that they had built their system to meet their needs. Google has now moved from the original Google File System and MapReduce to new implementations like Colossus (External link) and Dataflow (External link), as well as implementing their own relational database (External link).\nIn their critique of MapReduce, DeWitt and Stonebraker mentioned that MapReduce implementers would do well to study a bit of the history of parallel DBMS research, and it appears that Google did its homework. Moreover, the researchers mentioned that overall computer science communities tend to be insular and do not read the literature of other communities, and this urge to learn from other areas of computer science might be the biggest lesson to learn from MapReduce.\n","permalink":"https://nachooglpz.pages.dev/blog/003-what-happened-to-map-reduce/","summary":"David J. DeWitt and Michael Stonebraker argued that MapReduce was a step backwards for the database community.","title":"What Happened to MapReduce","topics":["database","dataflow-engines","Google"],"type":"essay"},{"content":"Rules, Sins, Virtues, Gods and more of The Church of EMACS according to The Gospel of Prophet Antony.\n","permalink":"https://nachooglpz.pages.dev/archives/002-gnu-gospel/","summary":"A Joke about the Church of EMACS.","title":"GNU Gospel","topics":["Open Source"],"type":"Web site"},{"content":"Alan Perlis once famously wrote (External link):\nIt is better to have 100 functions operate on one data structure than 10 functions on 10 data structures.\nA modern example of this aphorism can be found in object-oriented programming. From Perlis\u0026rsquo;s notion, you would rather write 100 functions for a single interface, than to write 10 functions for 10 different classes implementing that same interface.\nHaving the same interface and data structures to manipulate allows external functions and objects to apply the same protocols to the data structures they are working with; and the magic lies in the fact that a programmer can implement a use case for these interfaces (or objects) that the original programmers might never have thought of!\nA good example to apply this knowledge is in Lisp (and overall functional programming) (External link), as its basis lies in the manipulation of lists as a shared data structure to compute on. Having several functions that allow you to manipulate the same data structure allows you to combine and implement these functions in the way that you need, the same way as in the previous example applied to object-oriented programming.\nThe buffer type Probably one of the best implementation of an environment of functions (and plugins, or extensions) operating into the same data structure is Emacs (External link). The base protocol being implemented for these extensions is just the buffer type (External link), for Emacs Lisp (or elisp), the language to program any functions and extensions for the editor.\nOverall, the buffer is what makes Emacs the versatile, almost OS-like editor, that the Church of Emacs (External link) defends in the Editor war (External link) because the buffer is to emacs what the file is to UNIX as its fundamental data structure.\nAlthough when starting with Emacs the buffer can seem to be just a tab to represent a file, these are actually just objects that hold text that can be edited. The official documentation establish that most buffers do hold the contents of disk files to edit them, but because a buffer can represent anything (Alan Perlis also wrote that strings are the only communication coin we can count on), it can be used to manage other text-represented data (External link), such as binding network connections and external processes to buffers, or allowing asynchronous streaming of incoming data into buffers to be able to transform and manipulate them.\nAnother part of the magic for Emacs buffers is that, because they really are independent objects, you can have buffer-local variables, transaction queues for asynchronous processes, buffer-specific behaviors, and many other segregated functionalities implemented for each buffer.\nOperating into the same object What inspired me to write this entry was a vlog by Tsoding (External link), where he describes the usefullness of Emacs as annoying in the sense that its (useful and adaptive) design might make it difficult for any programmer to move to another editor. And this design lies in the buffer.\nAs an example, Tsoding demonstrated how he frequently uses two modes in Emacs: Dired (External link), and Multiple Cursors (External link). He mentioned that the developers for these programs developed them at completely different times and without designing them to be used together. Yet, they can indeed be used together without any friction at all.\nTsoding also mentioned that when reflecting on why could this happen, he thought of how they both were made to operate inside of a buffer instance, and this is why they can coexist and one can manipulate the other without any problem. And from his reflection, having the buffer type as the object target for programs is what makes Emacs the powerful tool that it is, as it allows programs to be able to be executed within each other to make a useful editor.\nAsynchronous processes Since Emacs allows you to run different programs to modify its buffers, the historical implementation that it has done to solve this problem is through asynchronous processes (External link). This allows Emacs to run its lisp programs without halting execution of the program, allowing users to continue modifying a buffer even if the running program has not finished its process. This asynchronity dates back to its first public release in 1985 (External link), and followed the trend of implementing event-driven design (External link) for Graphical User Interfaces.\nThe first Emacs (External link) was developed in 1976 for the Incompatible Timesharing System (ITS) (External link). One of the features of ITS was that a parent process had control of a child process, allowing it to halt it, resume it, and read or write its memory. And Richard Stallman had extended TECO (the predecesor to Emacs) with a \u0026ldquo;real time\u0026rdquo; full-screen mode, which would suggest that Emacs and its design was actually a precursor for event-driven design for applications.\nA developer on Hacker News (External link) described the Emacs buffer-based async handling of external processes as\nA masterpiece of software engineering, much better than things like UNIX pipes.\nWhile in UNIX pipes bytes flow from process A to process B, Emacs buffers processes allow for the asynchronous processes to coexist inside of the application while the programs remain easy to write under the same buffer data structure (and therefore programming under the same protocol).\nThreads Emacs is a product of its time, and even though it might have pioneered asynchronous implementation for its processes, it was made to run on a single CPU thread. This changed when in 2018 the editor added support for concurrency (External link), where each thread has its own current buffer and match data.\nTo maintain backwards compatibility with the existing implementation for Emacs, a process is locked to the thread that created it (External link), and the output from that process can only be accepted by that same thread.\nThere are several consequences of updating the editor to allow for use of threads from its initial architecutre, and some of the problems (External link) were noted by Tom Tromey (External link), one of the contributors for the existing threading code.\nTom noted that he tried the idea of having a per-buffer lock acquired by switching to the buffer. This idea did not work because, if you wanted to start a thread, you also have to make a buffer for it, or else it would try to lock the current buffer. Another problem with this per-buffer lock was that a process could not filter what was being written to the current buffer.\nTom also suggested that it would be possible to allevaite or remove the implementation that allows for only one Emacs Lisp thread to execute at a time. He ntoed that there was a patch once to alleviate it a bit, which would allow thread-switches periodically, instead of during I/O; but was rejected by the maintainers. He suggested that removing this global lock would be harder, and the main issue being that it would invole an audit of the core Emacs codebase to make sure there would not be any data races, or heap writes to be atomic.\nAnother user (External link) suggested using Transactional Memory (External link) to allow multiple threads to modify an arbitrary number of variables concurrently without locks, and in case of collision when merging the transactions, the thread working with stale data would restart its work. The noted problem (External link) was that this would need a stateless and restartable design implementation for the elisp programs for Emacs, which is not the case for the majority of existing programs.\nA different user on Hacker News (External link) noted that one challenge for implementation of threads on Emacs is that there are many elisp libraries that make assumptions about sequential and exlusive execution.\nTom had already concluded that the problems of adapting Emacs to thread usage made the threading work a failed experiment.\nEven though the design of Emacs might not be fit for supporting multi-threading, I would still argue that its pioneering desing focused on asynchronous processing of programs that would feed into a same buffer object, made this nearly OS-like editor the success it became; and the lightweight nature (External link) of Emacs justified not needing different processing threads from the start.\n","permalink":"https://nachooglpz.pages.dev/blog/002-the-power-and-limits-of-emacs-buffers/","summary":"Buffers are Emacs\u0026rsquo;s most powerful tool to be \u0026ldquo;annoyingly\u0026rdquo; useful.","title":"The Power (and Limits) of Emacs Buffers","topics":["open-source","design","data-structures","Emacs"],"type":"essay"},{"content":"One of the core sections of a Database Management System is the Buffer Manager (External link). It is responsible for managing which \u0026ldquo;pages\u0026rdquo; of data are stored in memory, and processing page requests that the File and Index Manager (External link) makes to fetch data.\nBecause a database already has a buffer manager, it might be better to actually bypass the operating system\u0026rsquo;s buffer cache as the policies that an operating system and a database system implement to store pages in their caches may differ. This is why a database system often also implements its own Disk Space Management (External link) module. This is the module that allows the buffer manager to retrieve pages from disk when the database needs to process them, or flush pages back to the storage disk when the buffer manager decides they are not needed anymore.\nBut to be able to bypass the operating system\u0026rsquo;s cache, and to be able to request data pages directly from the storage disk, a DBMS needs to fetch data using a different approach.\nReading data The first thing that a database does when working with data is to fetch it. To be able to make a read() call to the kernel without storing the page in the kernel\u0026rsquo;s cache, one can just use the O_DIRECT flag to enable direct I/O when opening the file.\nint fd = open(\u0026#34;file\u0026#34;, O_DIRECT); read(fd, buffer, 4096);It is important to note that O_DIRECT requires that buffer, offset, and I/O size align to the block size (External link), since the database is also bypassing the kernel\u0026rsquo;s ability to reshape the request into block-device-friendly operations; although this depends on the filesystem, kernel, and the storage device.\nNow the data is stored in a buffer to be able to manipulate it.\nWriting data Once the database has finished manipulating the buffer, or needs to flush the changes back to disk, one can simply use a write() call.\nwrite(fd, buffer, 4096);In a normal case, if the file was opened without the O_DIRECT flag, the write() call would return once the page was written to the kernel\u0026rsquo;s page cache, and the kernel may write it to storage later. But since we are using O_DIRECT, the I/O operation bypasses most of the kernel\u0026rsquo;s cache and goes directly to disk. It might also be important to remember that when doing a write() call, the kernel blocks this page until the call returns.\nShared buffers and fsync() With the previous approach, every time that a page is being written it needs to be flushed to disk bypassing the kernel\u0026rsquo;s cache, and this can be quite performance costly. Databases like PostgreSQL (External link) implement what they call shared_buffers.\nPostgreSQL actually accepts the trade-off of double buffering. This might have the advantage of keeping control of its own buffer cache to keep database semantics, while minimizing the performance penalty of a cache miss if the page that the database wants to retrieve still remains in the kernel\u0026rsquo;s page cache. PostgreSQL developers have designed its buffer management with the assumption that the OS cache is useful. Yet, a possible critique to this technique might be that when using a shared buffer, the database might be wasting space with some data pages being repeated in both buffer caches.\nThere are still operations where data pages still need to be written to disk just after they happen in the database, like writes to the Write Ahead Log (External link) (or WAL) which might be useful for database transaction processing. The way that they ask the kernel to flush these pages is through the fsync() call.\nThe database might be performing a series of writes, which will be stored in the kernel cache, and then the database will ask the kernel to flush the buffer back to disk.\nint fd = open(\u0026#34;file\u0026#34;, O_RDWR); write(fd, buffer, 4096); write(fd, buffer, 4096); write(fd, buffer, 4096); fsync(fd);Other embedded databases also rely on buffered I/O and kernel page cache.\nPersistence Everything has probably seemed straightforward up to this point, but the pièce de résistance lies in the persistence of the write() operation that the database wants to make.\nMaking a direct write, or using fsync() to flush a page to disk, does not mean that the write will persist on disk. Phil Eaton (External link) once made a list of Things that go wrong with disk IO (External link).\nThe ones concerning the aforementioned write mechanisms are the following.\nData got corrupted\nData might corrupt at any moment during the journey from the computer\u0026rsquo;s memory to disk. One defense against silent corruption is to store checksums alongside data and verify them when the data is read. Databases like PostgreSQL before version 18 (PostgreSQL 18 now does) and SQLite do not checksum, though. Another caveat is that, once a corruption has been detected, recovery requires another source of data, such as WAL, a replica, redundant storage, etc.\nData was partially written\nWhen a page arrives to disk it is possible that some sectors of the disk where that page belongs get written, and then the system crashes without the page being fully written on disk.\nThis is called a torn write. A way to overcome this issue might be to duplicate all the writes, although it is important to note the tradeoff that a twice-as-costly write would represent. Other ways (External link) include logging page deltas, Copy-On-Write B-Trees or Copy on First Write.\nData did not reach disk\nMost drives contain a volatile write cache to speed up sequential writes (External link), and because the write operations return once the file descriptor is transferred to the hardware device, the kernel washes its hands after it delivers the page to the hardware and forgets about it. The hardware might store this write on its own cache and, in case of a hardware crash, the write might be lost; although it is important to note that modern standards (External link) for persistent writes, like Force Unit Access (FUA), allow to guarantee a write once the page arrives to the hardware.\nfsync() failed\nfsync() does not guarantee to succeed, and when it fails, it reports a failure to all file descriptions that were open at the time of failure, and the only way to know if your write did actually fail is to read again the page. The only way to know which exact write failed is through O_DIRECT.\nFsyncgate When fsync() fails, the kernel returns the error code and considers its job done: it is now the program\u0026rsquo;s problem to recover. Before 2018, PostgreSQL was hoping that whenever fsync() failed, the kernel would remember the partial write and try again later (External link) until an fsync() would work and store the data on disk. PostgreSQL assumed that the dirty data would remain available for a later synchronization attempt.\nThere were some problems with this approach. The first one is that if the write to disk was failing, you still had to store the data somewhere, so the kernel and database buffers would fill up. Second was the implementation that was made using this API: when two processes tried to write to the same file, but the first process did not flush with fsync(), the kernel scheduled the buffer\u0026rsquo;s flush to disk at a later time. But then a second process would submit a write (again, without fsync()) and the kernel decided that it was time to flush this write to disk, but discovered that the write did not work, and reported the failure to the second process. The first process decided to flush its changes to disk using fsync(), and because there were no pending writes, the operation would succeed.\nFrom the kernel\u0026rsquo;s point of view, this was consistent: it reported the error when the first flush failed, and when the fsync() call was made, since there were no pending writes, there was nothing to do, and did not return an error. But from an application\u0026rsquo;s point of view, this might not have been the optimal behavior, as the first process did a write and an fsync(), and got no error.\nThis led to a user finding data corruption after a storage error (External link), and what would be later known as \u0026lsquo;Fsyncgate\u0026rsquo;.\nCompletion semantics When using O_DIRECT, the kernel avoids copying data from user space into kernel space, and it instead writes it directly via DMA (Direct Memory Access). But there is no guarantee that the call will return only after all data has been transferred, so the kernel may see the write operation finished before the data is physically written to disk.\nI previously mentioned that there are modern standards for persistent writes, and one of them is what we know as Forced Unit Access (FUA) (External link). When we enforce FUA on a write() command, the operation will return once the buffer is written to disk (and not the disk\u0026rsquo;s cache).\nThere are two flags (External link) that enable completion semantics:\nO_SYNC guarantees that the contents of the file are written to disk. O_DSYNC guarantees not only that the file has been written to disk, but its metadata as well. It is important to note that using either of these flags does not mean that the program is bypassing the disk\u0026rsquo;s cache, but instead, the hardware is guaranteeing that data being written will persists. The mechanism to do so depends on the hardware, as there are HDDs and SSDs that implement non-volatile cache that ensure that even in a crash or power failure, data still manages to be flushed inside of the hardware.\nAnother approach to O_DIRECT During 2002, Linus Torvalds had proposed another implementation for O_DIRECT (External link). During that time, O_DIRECT was not as performant, and it showed up to a 55% performance hit vs no O_DIRECT. Linus attributed this problem to the fact that O_DIRECT needed to be asynchronous and had to do read-ahead.\nWhat Linus had proposed when doing an O_DIRECT read into a buffer is to divide this process in two phases:\nAllocate the pages, and start the I/O operation asynchronously. mmap the file with a MAP_UNCACHED flag, causing read-faults to \u0026ldquo;steal\u0026rdquo; the page from the page cache and making it private to the mapping on the page faults. And any write() operation would be the other way around:\nTake the pages in the memory area and move them to the page cache, removing the page from the page table (and only copying it if pages already exist). Make the I/O operation to disk. With this approach, the kernel would not have to make a copy of the buffer that the process is trying to flush to disk into its own cache (to immediately send to disk), but would just take ownership of this allocated buffer from the process, and be able to do any operations it needs with it (including flushing it to disk).\nWhile a clever implementation that would make O_DIRECT operations more performant, the only change that was made to the design was the asynchronous processing of O_DIRECT. Because most of the database systems (and other complicated programs) were already using the POSIX-inspired API, the only update that was made to the design of this API was the underlying processing for these calls.\n","permalink":"https://nachooglpz.pages.dev/blog/001-how-databases-write-to-disk/","summary":"A successful write by a database does not necessarily mean durability on disk.","title":"How Databases Write to Disk","topics":["databases","syscalls","i/o"],"type":"essay"},{"content":"During 2002, Linus Torvalds had proposed another implementation for O_DIRECT (External link). During that time, O_DIRECT was not as performant, and it showed up to a 55% performance hit vs no O_DIRECT. Linus attributed this problem to the fact that O_DIRECT needed to be asynchronous and had to do read-ahead.\nWhat Linus had proposed when doing an O_DIRECT read into a buffer is to divide this process in two phases:\nAllocate the pages, and start the I/O operation asynchronously. mmap the file with a MAP_UNCACHED flag, causing read-faults to \u0026ldquo;steal\u0026rdquo; the page from the page cache and making it private to the mapping on the page faults. And any write() operation would be the other way around:\nTake the pages in the memory area and move them to the page cache, removing the page from the page table (and only copying it if pages already exist). Make the I/O operation to disk. With this approach, the kernel would not have to make a copy of the buffer that the process is trying to flush to disk into its own cache (to immediately send to disk), but would just take ownership of this allocated buffer from the process, and be able to do any operations it needs with it (including flushing it to disk).\nWhile a clever implementation that would make O_DIRECT operations more performant, the only change that was made to the design was the asynchronous processing of O_DIRECT. Because most of the database systems (and other complicated programs) were already using the POSIX-inspired API, the only update that was made to the design of this API was the underlying processing for these calls.\n","permalink":"https://nachooglpz.pages.dev/archives/001-o_direct/","summary":"Mail discussion about Linus Torvald\u0026rsquo;s redesign proposal for O_DIRECT.","title":"Linus's Proposal for O_DIRECT Redesign","topics":["Open Source","Linux"],"type":"Mail"}]