[{"content":"After playing with Windows XP on my \u0026ldquo;new\u0026rdquo; Optiplex 990 I asked myself a question: \u0026ldquo;Would Windows 98 SE run on this Sandy Bridge as well?\u0026rdquo; By running I mean: accelerated 3D graphics and 3D audio (EAX) without any hangs or lockups, being able to play games and run benchmarks.\nComponent my Optiplex 990 configuration CPU Core i5 2500 Chipset Q67 (Sandy Bridge, 6-series PCH) RAM 4GB, 2+2, dual channel DDR3-1333 GPU Radeon X550, 64bit (most retro part in here) Sound Sound Blaster Audigy (SB0090) in the PCI slot Storage 128GB SATA (BIOS SATA == AHCI) Other USB keyboard and mouse, one PS/2 mouse so I could install USB drivers Turned OFF: onboard: GPU, NIC, sound; UEFI boot (legacy boot), all power saving/sleep states, virtualization, remote management Table 1 — the test subject.\nI knew it wouldn\u0026rsquo;t be easy.\nThere is a reward though. Apart from bragging rights, Win98 is an incredibly lightweight system. We are talking about megabytes here - whether it is the RAM used or files loaded at startup. It feels like firmware today and loads instantly. It is still DirectX 9.0c capable, though, and that is something to think about.\nThe Journey This whole journey is enticing because detailed control over the hardware must be in your hands - Windows 98 does not know what surrounds it in most cases.\nGoing through the exact process of exploration would be pointless. It took me over two weeks of my free time and I went through ups and downs, overcoming one obstacle after another, only to finally realize that I had met the goal. I felt the joy of an achievement which only a true engineer could appreciate.\nI skipped all theoretical preparation: loaded Easy2Boot onto a USB key, ran FreeDOS fdisk and partitioned my 128GB SSD. A trusted Win98 SE OEM ISO came next, it can boot by itself, a great feature for a machine without an FDD/emulator.\nHIMEM.SYS refused to load, so there was no extended memory to boot with, and setup.exe hung after every workaround I tried. It was official: The experience from installing Win98 on Core 2 Duo systems won\u0026rsquo;t transfer to a Sandy Bridge platform.\nI started to treat this project seriously, with respect. As if it were a paid project I had to debug.\nThe Lab I established the following facilities for my diagnosis and testing efforts:\nDOSBox-X and QEMU virtual machines on a separate PC Debian 13 on a secondary SSD (right there, in the Optiplex 990) a USB stick to carry all patches and files needed SATA-to-USB 3.0 adapter, so I could inspect logs and state on the Win98 partition between each attempt The SSD preparation was done completely from Linux (Debian 13) - cleaner and faster than tinkering with fdisk and period-correct tools. I copied all install files over and also system files from a Win98 boot floppy, modern components (XMGR, patched CAB, W9xFix) and tools like EDIT or Volkov Commander to make it more convenient to adjust files in DOS mode.\nThe setup was started with setup.exe /IS /IE /IM /NM /p i. That skips the startup disk wizard, memory and machine checks, ScanDisk and most importantly /p i tells Setup not to report the Plug and Play BIOS at all, resulting in a \u0026ldquo;standard PC\u0026rdquo; installation. The board is ACPI of course, but its tables are 12 years younger than the parser and would just confuse Win98 unnecessarily. Linux later proved that ACPI mode would not have helped anyway. The ACPI tables route the Management Engine (MEI) onto a shared line just the same.\nThe Intel Management Engine is an autonomous subsystem that has been incorporated in virtually all of Intel\u0026rsquo;s processor chipsets since 2008. The Intel Management Engine has been renamed to Intel CSME - Intel Converged Security and Management Engine (several years after Optiplex 990 was released). Wikipedia\nThe setup was pre-patched using WIN98_54.CAB, the cabinet with rloew\u0026rsquo;s RAM Limitation Patch already applied to VMM and VCACHE. Setup installs the patched kernel from the first GUI start, so no 1GB limit at any point.\nGetting Setup through Initial attempts were hanging right at the start. I had to focus on the HIMEM and setup.exe hangs.\nOne obvious problem was the A20 gate, which is by default toggled by HIMEM on startup. It cannot be turned off on this PC. HIMEM can leave it alone with the /a20control: switch. Set /A20CONTROL:OFF and HIMEM will only take control of the A20 line if it was turned off when the driver loaded. If the A20 line is already active, HIMEM leaves it alone. After that it hung anyway. I didn\u0026rsquo;t investigate further and switched to XMGR.\nXMGR is a DOS extended memory manager (XMGR.SYS) that handles up to 4GB of XMS (Extended Memory Specification) RAM. Written by Jack R. Ellis, it serves as a lightweight alternative to traditional managers like HIMEM.SYS. Vogons Wiki\nThe second hang was right at the start of the GUI part of Setup. The IOS.VXD hang failed to reproduce in emulators, but finally the disassembly provided a hint: It reads sector 0 and if the 16-bit word at offset 0xDA and the 32-bit dword at offset 0xDC are both zero, it builds a \u0026ldquo;disk stamp\u0026rdquo; from the BIOS time, writes the MBR back with INT 13h, re-reads it and compares. By trying this operation in isolation I verified that a CHS write like that never returns from the BIOS on this PC. My workaround was simply to pre-stamp the drive from Linux so IOS.VXD never enters this path. The problem technically remains unsolved.\nIOS.VXD is the core 32-bit I/O supervisor virtual device driver in Windows 9x and Windows Me operating systems.\nThe next obstacle was a device detection hang at 33% of phase 2, after the first reboot once all files are copied. This turned out to be a parallel port which is not even physically present on this PC. The board PCB has some space for a header, but nothing is soldered on and the BIOS is completely oblivious about it, not even offering an option to disable the parallel port. The fix was easy, and I must credit the original Win98 setup here: reboot. After the hang, the setup blacklists the crashed probe and never looks for the device again.\nThe OS boots up. It is very slow and unstable, every filesystem operation causes a freeze, peripherals are working as emulated PS/2 devices. There is no sound, no GPU acceleration, just a 16-color 640x480 desktop.\nGetting it stable Initially, I suspected some clashes of devices on IRQ lines, so I tried to separate and shuffle them around. This part proved to have multiple root causes ultimately. Some were uncovered thanks to blue-screens and the rest of them only thanks to having a secondary OS at hand. Using alternative boot profiles I was able to look \u0026ldquo;through the eyes\u0026rdquo; of Win98 with better tooling. Booting with acpi=off noapic (the same legacy PIC world Win98 lives in) showed IRQ 11: nobody cared, 100k interrupts in a few seconds. Handlers on the line: pcieport, ehci, firewire, mei_me. Once mei_me was blacklisted the count dropped to 31.\nMultiple devices I didn\u0026rsquo;t care about were active, but driverless (in Win98), causing interrupt storms. The most dangerous was the MEI, which could not be disabled in this BIOS. I implemented a tool which (on every boot) sets the INTx-disable bit for every such device (function). W9xFix The EHCI hand-off proved to be a problem. The NUSB package correctly installs drivers for all USB controllers and devices, but NUSB does not do the EHCI BIOS-to-OS hand-off. The PS/2 keyboard and mouse emulation (BIOS SMM handler) could not be controlled in the BIOS. It is always ON, even after the USB driver takes over. Windows boots after about 100s, and then the mouse and keyboard update about once per second. Obviously, the stack works only on timeouts, not as intended. This is another fix done within my tool, right before booting Windows: clear the SMI enables, set the OS-owned bit, wait for the BIOS-owned bit to drop. w9xfix ehci handoff. The hand-off kills the BIOS keyboard emulation, so it must be the last step before Win98 starts and only after NUSB is installed. The last but critical issue for the system stability is the ASPM mismatch. The BIOS enables ASPM on old GPUs like my Radeon X550. The GPU\u0026rsquo;s acceptable exit latencies are below what the root port offers, so by the PCIe rule neither L0s nor L1 may be enabled, yet the BIOS enables both. This was not apparent at first because the GPU was working normally in 3D, even able to run a whole 3DMark test (or two), only to suddenly freeze mid-game. After I crosschecked the GPU under load in Linux, knowing it was not faulty, I spotted the mismatch. This is the third and last fix provided by my tool: w9xfix aspm links 0 walks every root/downstream port, every function behind it, clears ASPM on the device end first and then the port. There was one hidden (now obvious) land mine: the Radeon X550 (and ATI cards from that era in general) appear as two devices (the second display head is its own PCI function). The ASPM has to be set for both. This is why my early attempts at this fix were failing. The AHCI driver clashed with the sound driver. This is an interesting coincidence, but two third-party VxDs picked the same ID from the \u0026ldquo;unofficial\u0026rdquo; range. Namely it was the Sound Blaster driver (emu10kx VxD) clashing with rloew\u0026rsquo;s native AHCI patch/driver. The fix is to patch the device ID to 0x0000 (UNDEFINED_DEVICE_ID) in both places AHCI.PDR carries it, the LE header and the DDB. This is fine since Windows matches the controller through the INF on its own. The three fixes W9xFix applies on every boot, and how portable they are:\ncommand how it is done where it works intx ... off sets bit 10 (INTx disable) in the PCI command register of the selected functions any PCI 2.3+ device (2002 and later); older devices ignore it aspm links 0 walks the PCIe capability of every port and every function behind it; clears the device end first, then the port any PCIe system ehci handoff writes USBLEGSUP and USBLEGCTLSTS in config space at the assumed EECP 68h, capability ID checked (eecp=XX overrides) Intel ICH/PCH and most others; moot on Skylake and newer, which have no EHCI Table 2 — W9xFix boot-time fixes and their portability.\nThe Result Setup was then finalized using standard VxD drivers for the SB0090 (SB Audigy) sound card, the Catalyst 6.2 driver for the Radeon X550 GPU and the DirectX 9.0c runtime.\nThe paradox of the AHCI driver under Windows 98 on this board: Technically it is easier to get working than the alternative, legacy IDE emulation in the BIOS, and it is naturally faster. What underlines the paradox: I failed to get any AHCI driver, official or not, working under Windows XP on the same machine. There it runs on emulated IDE only.\nI have a fully working and very snappy Windows 98 installation now.\nThis story is also an episode on YouTube. I am covering additional experience with this setup in my next article. Watch the optiplex-990 topic and the YouTube channel for more details.\nReferences VasilijP, W9xFix, the boot-time PCI fix tool built for this project. Jack R. Ellis, XMGR, DOS extended memory manager, FreeDOS help. \u0026ldquo;Intel Management Engine\u0026rdquo;, Wikipedia. \u0026ldquo;Architecture of Windows 9x\u0026rdquo;, Wikipedia. \u0026ldquo;Useful DOS utilities\u0026rdquo;, Vogons Wiki. ","date":"2026-09-12T00:00:00+02:00","image":"/p/win98-se-on-sandy-bridge-debugging/cover_hu_10be7450723edef1.png","permalink":"/p/win98-se-on-sandy-bridge-debugging/","title":"Windows 98 SE on Sandy Bridge: the debugging story"},{"content":"Some of the games I loved on my 386 never got a source release. The companies moved on, the toolchains died, and what remains is a couple of hundred kilobytes of 16-bit machine code on a floppy image. For years I assumed that really understanding one of these games - not just running it in DOSBox, but opening it up, seeing how it ticks, and maybe even improving it - was a job for someone with infinite patience and a decade of spare time.\nI no longer believe that. Over the last few months I have been digging into a 1991 DOS flight simulator (Chuck Yeager\u0026rsquo;s Air Combat, a childhood favorite) as a long-running side project.\nSome approaches work far better than I expected, some are dead ends, and the difference is worth writing down. My tools: C#/.Net and my knowledge of how to push pixels onto a screen using the CPU. In this round, AI was also a big help when it came to the tedious work of analyzing a binary executable piece by piece.\nWhat works: your own CPU emulator, in-process The first surprise: emulating an old real-mode x86 CPU is not a heroic undertaking. It is a very approachable piece of engineering, and having your own emulator - in your own language, in-process, under your debugger - changes everything.\nThe whole machine state of a real-mode PC is charmingly small:\nMemory is an array of bytes. One megabyte, flat, new byte[0x100000]. Segmented addressing (segment * 16 + offset) looks scary in old books but is one line of code. The register file fits in a screenful. AX/BX/CX/DX, the index and stack registers, segment registers, FLAGS, IP. Done. The instruction set is finite and mostly regular. A fetch-decode-execute loop plus a few hundred instruction handlers gets you to \u0026ldquo;the game boots\u0026rdquo;. The 8086 has no caches, no pipelines, no protection rings to model - every instruction is just a small pure function over that byte array. // The heart of the whole thing is not much more than this: while (running) { byte opcode = mem[cs * 16 + ip++]; switch (opcode) { case 0x8B: /* mov r16, r/m16 */ ... case 0xE8: /* call rel16 */ ... // ... a few hundred friends ... } if (pendingInterrupt \u0026amp;\u0026amp; interruptsEnabled) DeliverInterrupt(); } Around the CPU you need a small cast of devices:\nThe timer (PIT). A counter that raises IRQ0 at a programmable rate. Games of the era live and die by this interrupt - it drives their frame pacing and their sense of time. The keyboard controller. Scancodes in a queue, IRQ1, port 0x60. The VGA adapter. This is the biggest device, but it is well documented: a 256-entry palette DAC, and for this game an unchained, planar variant of the classic Mode 13h - 320×200 with four bit-planes, where the Sequencer\u0026rsquo;s Map Mask register selects which planes a write lands in, and pixel x lives in plane x \u0026amp; 3. Emulating it means modeling a handful of I/O ports and the planar memory layout; rendering means folding four planes back into one linear image. DOS itself - mostly INT 21h file I/O - can be faked at the interrupt boundary with a few dozen functions. No real DOS needed. None of this is research; all of it is documented (FreeVGA, the Intel manuals, Ralf Brown\u0026rsquo;s interrupt list). The reward for doing it yourself instead of using an off-the-shelf emulator is enormous: every byte of memory, every port write, every interrupt delivery is your object model, one function call away from your analysis code. You can put a callback on \u0026ldquo;any write into video memory\u0026rdquo; and ask which code wrote this pixel? This single question is worth the entire emulator.\nThe picture above shows my Avalonia UI running the emulator (which is a CLI tool). This is what my emulator looks like after 3 months of tinkering.\nA modern CPU also makes brute force respectable: my plain C# interpreter replays about 30 million emulated instructions per second single-threaded. A full 15-minute gameplay session - several billion instructions - replays in a couple of minutes.\nWhat works even better: deterministic recording The second pillar, and honestly the load-bearing one: make the emulation perfectly deterministic, then record sessions.\nOld games are nearly deterministic already. The nondeterminism comes from a short, findable list of sources:\nthe timer interrupt arriving \u0026ldquo;whenever\u0026rdquo;, relative to the executing code; keyboard/mouse input arriving \u0026ldquo;whenever\u0026rdquo;; the RNG being seeded from the wall clock or timer phase; anything that reads a hardware counter mid-computation. So you hunt these down and pin them. Interrupts are not delivered \u0026ldquo;whenever\u0026rdquo; - they are delivered at an exact, counted instruction boundary. Input events are not injected in real time - they are stamped with the exact instruction count at which they fire. Once every source of variation is indexed against the instruction counter, a gameplay session becomes a small file: the initial state plus a list of (instruction_index, event) pairs.\nReplaying that file reproduces the run bit for bit. Same billions of instructions, same memory image at the end, same final framebuffer hash. Every time, on any host.\nAnd here is the beautiful part: playing the game becomes writing a test suite. Fly a mission, eject, watch the parachute from the external camera, wander through the menus - save the session, and you have a regression test that exercises exactly those code paths forever after. My test battery is currently a set of such recordings (now 159 and counting); any change to the emulator (or to the lifted code, below) must reproduce every recording\u0026rsquo;s final state hash exactly. A single wrong flag bit in one instruction handler shows up as a diverged hash within seconds. It is the strongest safety net I have ever had in any project, and it costs nothing but a few bytes of disk space.\nThe strategy on top: \u0026ldquo;lifting\u0026rdquo;, supported by decompilation With a deterministic emulator and a recording battery, you can do something that would otherwise be reckless: replace pieces of the original machine code with native, readable code - one function at a time - and prove each replacement correct.\nThe workflow looks like this:\nProfile. Count executed bytes per function across the recordings; sort. The hot spots are always fewer than you fear - polygon fillers, line drawers, the flight model integrator, the mission state machine. Decode. Disassemble the function; use a decompiler (Ghidra) as a map, but treat the bytes as the ground truth. Write down its exact contract: every register it reads and writes, every global, every byte of stack it touches, its exact instruction count for each path. Lift. Implement the same effect in C#, installed as a trap at the function\u0026rsquo;s entry inside the emulator. Shadow. For every call during replay, run both: the genuine 8086 code and the C# prediction, and compare the complete effect - every written byte, every output register, even the \u0026ldquo;garbage\u0026rdquo; the routine leaves below the stack pointer. Millions of live calls with zero mismatches, on real gameplay data. Certify. Finally, replay the whole battery with the lift active and demand the end-to-end result stays bit-identical to the pristine run. The emulator hosts the game; the game gradually becomes native code; and at every step the original binary itself referees the correctness of your understanding. There is no \u0026ldquo;I think this is what it does\u0026rdquo; - either the hash matches or it doesn\u0026rsquo;t. A couple of hundred functions in, the harness has caught wrong assumptions I would never have found by staring at disassembly: an undocumented flag quirk, a routine whose \u0026ldquo;obvious\u0026rdquo; output register was actually dead, a comparison whose equality case short-circuits differently.\nThe game\u0026rsquo;s executable was triple-packed - a linker-level compressor on top of an EXEPACK-style layer on top of the actual code. My first instinct was to identify the packers and reimplement the decompression. The better move turned out to be embarrassingly direct: run the unpacker stubs in the emulator and dump the memory image after they finish. The packer authors already wrote a perfect decompressor; it was sitting right there in the first kilobytes of the file. An afternoon of work instead of a week of format archaeology.\nWhat fights back: friction points Not everything is smooth sailing. A catalogue of the things that actually cost me time:\nInterrupts versus lifting. When you replace a 5,000-instruction span with one native call, the timer interrupt that would have arrived in the middle of that span now has nowhere to land. You can defer it to the end of the span - but this game samples its tick counter at one precise spot in the frame loop, and a tick delivered on the wrong side of that read changes the frame\u0026rsquo;s delta-time and forks the entire subsequent trajectory. Getting a policy for this that provably preserves behavior (and knowing when it can\u0026rsquo;t) was the single hardest design problem of the project - much harder than any individual function.\nThe hand-written assembly of the era. Compiler-generated code is pleasantly boring. But the hot paths were written in assembly, and those developers did not respect anyone\u0026rsquo;s calling convention: functions passing arguments in whatever registers were handy, two languages\u0026rsquo; conventions (C and Pascal) mixed in one binary with opposite argument orders, routines that mutate their caller\u0026rsquo;s stack slots.\nSelf-modifying code (SMC). The renderer patches instruction operands in place - a color byte here, a jump displacement there - as its normal mode of operation. One step further, the projection routine is generated at runtime into a buffer, parameterized by the current zoom scale. You cannot \u0026ldquo;just disassemble\u0026rdquo; code that doesn\u0026rsquo;t exist until the game is running. (The emulator saves you again: watch writes into code regions, catch the generator in the act, and prove the generated code is a closed family of templates.)\nExecutable code hiding in data. A third of the video-memory writes in a session came from code that is not in the executable at all - it lives compressed inside the asset archives and is loaded like any other resource. Mission files carry little x86 routines implementing their victory conditions. Sound drivers are loadable modules. The \u0026ldquo;binary\u0026rdquo; you must understand is scattered across the whole game data.\nThe math of a machine with no FPU. There is not a single floating-point instruction in the image. Everything is fixed-point: angles as binary fractions of a circle, slopes in Q14, sine via lookup tables (with off-by-one sized tables and out-of-range reads that are bounded-wrong on the 8086 but crash a naive port). Multiplication was expensive, so you find shift-add sequences and precomputed tables everywhere; division was dangerous, so you find guards that do double duty - my favorite: the flight model\u0026rsquo;s ±80° pitch clamp turns out to be the divide-by-zero protection for the trigonometry behind it. Misread one of these tricks and your port is subtly, maddeningly wrong.\nThe payoff: before and after Why go through all of this instead of just enjoying the game in DOSBox? Because once the emulator understands the renderer well enough to lift it, you can do more than reproduce it - you can refine it. The lifted drawing functions see the game\u0026rsquo;s drawing commands above the 320×200 quantization: polygon vertices, spans, circles, glyphs, before they are crushed onto the coarse grid. Feed those same commands to a modern software rasterizer - true-color, anti-aliased, at 4× the resolution - and you get the original scene, the original geometry, the original palette\u0026hellip; just cleaner.\nThe same scene, seconds apart, switching renderers live in the emulator:\nBefore - the original renderer (320×200, upscaled):\nAfter - the refined renderer (same geometry, rendered host-side at 1280×960: 4× the resolution at 1:1 pixels, aspect-corrected for the 4:3 screen the original 320×200 mode was physically displayed on):\nBefore/after pair of screenshots of the same scene: once through the original renderer, once through a \u0026ldquo;refined\u0026rdquo; renderer that draws the original game\u0026rsquo;s geometry at high resolution - the payoff of the whole exercise.\nLook at the tail fin, the canopy frame, the landing gear struts, the shadow under the fuselage. It is the picture the 1991 renderer was always describing, finally drawn with the pixels it never had.\nThe refined renderer in flight, 1920×1080:\nWhy did this become feasible? The ingredients that make it feasible today, in order of importance:\nMassive help of AI: I used Claude Code with the latest models. What would have taken years before now takes weeks or months (realistically, considering the token consumption limits). CPU cycles are free. Emulating a 16 MHz machine at thousands of times real speed makes replay-everything, verify-everything workflows practical. Determinism turns gameplay into tests. This is the idea I\u0026rsquo;d keep above all others - remove the randomness, index every input, and the game verifies your understanding of it, continuously. In-process beats off-the-shelf. An emulator you own is an analysis instrument, not just a player. The old tricks are learnable. SMC, fixed-point math, mixed conventions - each one is a speed bump, not a wall, once you can watch the code run instruction by instruction. The hidden treasures are still there in these old binaries - dormant features, cut content, elegant tricks nobody has seen in thirty years. The tools to dig them out fit in a hobby project now. If there is a game from your past whose source was never released: it is more within reach than you think.\nReferences Chuck Yeager\u0026rsquo;s Air Combat, Electronic Arts, 1991 — the game under the microscope. FreeVGA project documentation — the VGA hardware reference used for the adapter model. Intel 80386 Programmer\u0026rsquo;s Reference Manual — CPU semantics for the emulator core. Ralf Brown\u0026rsquo;s Interrupt List — the DOS and BIOS interrupt boundary (INT 21h and friends). ","date":"2026-08-03T00:00:00+02:00","image":"/p/old-school-graphics-part-3-hidden-treasures-of-old-games/banner_hu_941c198bc69d7b97.png","permalink":"/p/old-school-graphics-part-3-hidden-treasures-of-old-games/","title":"Old-School Graphics in C# / .Net 10, Part 3: Hidden Treasures of Old Games"},{"content":"A continuation of my previous article, What is better than your AI loop?, which introduced the blackboard architecture this post builds on.\nThe chat. The message-response chain you produce by typing with an LLM. It could be trivially implemented by a single growing context or a sliding window.\nWhy does it feel like an anti-pattern to me? Chat as a UI is fine, chat as a world model is not.\nSpeech Act Theory, pioneered by J.L. Austin, tells us that language is a form of action. In his 1955 lectures, published posthumously in 1962 as How to Do Things with Words [1], Austin dismantled the idea that sentences merely describe reality, introducing \u0026ldquo;performatives\u0026rdquo; - expressions changing the world state. (Coincidentally, 1962 is also the year the first \u0026ldquo;blackboard\u0026rdquo; appeared in the AI literature [2].) Later on, John Searle systematized the field, sorting speech acts into five distinct types (Representatives, Directives, Commissives, Expressives, Declarations) [3], and this taxonomy transitioned from philosophy to computer science and AI.\nWith the emergence of Multi-Agent Systems (MAS) in the 1990s, the traditional RPC style of communication became a bottleneck. The request-response was an impedance mismatch. Independent software programs needed to negotiate, plan, share knowledge. Autonomous agents required loose coupling simply to remain autonomous. An agent must be able to refuse a command, negotiate a task or report status on its own.\nTo standardize this for machines, the early 1990s brought KQML (as part of the DARPA Knowledge Sharing Effort) [4]. The Foundation for Intelligent Physical Agents (FIPA) released the FIPA-ACL specification in 1997 [5]. They grounded agent communication in Belief-Desire-Intention (BDI) modal logic (logic extended by necessity and possibility statements/operators). This brought problematic preconditions, for example: to use Inform, the agent must believe the proposition and not believe the receiver already knows it. That is unverifiable and the equivalent of mind reading. This rigid codification in FIPA-ACL was ultimately a hard problem, which is now elegantly (and perhaps unintentionally) approximated/solved by LLMs for free.\nThe formal schema wasn\u0026rsquo;t widely adopted by today\u0026rsquo;s LLM agents. Modern multi-agent frameworks (such as AutoGen [6] or LangGraph) largely abandoned strict modal logic. Instead they drifted towards natural language schemas and JSON wrappers (e.g. the OpenAI function-calling schema introduced in 2023). However, the foundational principles of Speech Act Theory remain present and identifiable.\nChat with Images Let\u0026rsquo;s consider a standard chat interface where a user uploads an image or possibly even multiple images across the conversation.\nI generated this test image in multiple variations and it is surprisingly tricky for most models to get everything right.\nIn a standard LLM chat app, the system relies on a sliding context window. To answer the second question, the model must either rely on a highly compressed, emergent internal representation of the image from the first pass, or it must re-encode the entire image alongside the growing chat history. It is treating the conversation as a \u0026ldquo;next token prediction\u0026rdquo; problem, while hoping that critical image detail was captured in its internal attention weights.\nThe same conversation in my blackboard UI - note the journaled image \u0026ldquo;Q\u0026amp;A\u0026rdquo;.\nSpeech Acts on a Blackboard A coding agent treats the chat (operator messages), its own responses and tool outputs as an event stream from which a BDI is parsed and used to derive its internal state, which is (usually) explicitly represented. The anti-pattern part here, which still remains, is the limited context window which accumulates the chat transcript. That would be a redundant non-issue in an ideal case, where all incoming inputs are fully represented in the internal state, which is unfortunately not the case.\nToday\u0026rsquo;s coding agents/harnesses are therefore only boosting the emergent model capabilities when it comes to world representation, not replacing it. They use tools, a large context window and clever internal logic to hide the inherent limitation. Ever had an agent forget something it was asked to do?\nOnce we map the chat or agentic use case to an event-driven, CQRS architecture, we are no longer dealing with RPCs masquerading as agent performatives. The performatives act as state mutations of the world model itself. The Vision Reader KS asks the image targeted questions, and the answers become durable evidence on the board. No re-encoding, no context truncation, no attention decay.\nThe diagram above shows a board projection using just 2 knowledge sources: Vision Reader KS (priority 90) and Chat Responder (priority 80). Blue nodes are operator messages, green nodes are placed by the Chat Responder. Beige is the image evidence and purple are question/answer pairs related to the image.\nThe diagram above adds an additional KS: Intent Spotter (priority 95). Intent nodes are colored light red. It is no longer the Chat Responder\u0026rsquo;s task to reason about the \u0026ldquo;what should be done?\u0026rdquo; part. Chat Responder simply takes the evidence produced on the board and formulates a response for the user.\nThe roster maps naturally onto both the classic cognitive-role split (Recognition → Planning → Execution → Reflection) and Boyd\u0026rsquo;s OODA loop - Observe, Orient, Decide, Act [7]:\nRole OODA KS in this article Speech act Board evidence Recognition Observe Vision Reader (priority 90) Representatives Facts, image answers. Recognition Orient Intent Spotter (priority 95) Directives Intent node extracted from natural language. Planning Decide Planner (Chat Responder, merged with execution in this case) Commissives A commitment to act, journaled before any action. Execution Act Chat Responder (priority 80) Representatives, Declarations The action performed exactly once, result recorded. Reflection returns to \u0026ldquo;Observe\u0026rdquo; done by the user Directives Difference between intended and achieved state, posted as a fresh intent. The hidden discrepancy between chat models and blackboard systems becomes obvious when we look at state persistence. Anyone who has used an LLM for complex coding tasks has experienced the model \u0026ldquo;forgetting\u0026rdquo; a constraint or an instruction from several prompts ago. This happens because the model\u0026rsquo;s intent recognition is tied to its attention mechanism. The attention of each LLM is diluted with growing context (the \u0026ldquo;lost in the middle\u0026rdquo; phenomenon [8]).\nIn a Blackboard CQRS architecture, an intent is an immutable event in a log. When a Knowledge Source posts a Directive or Commissive to the board, it remains there as an unresolved node. It cannot be forgotten by a sliding window. It sits in the event store until any KS posts a corresponding result that satisfies it.\nIf execution fails, there could be a \u0026ldquo;Reflection\u0026rdquo; KS to close the OODA loop. It observes the discrepancy between intended and achieved state and could post a new intent to simply try again. The LLM is relegated to what it does best: acting as the cognitive engine for individual Knowledge Sources, rather than being forced to act as the entire \u0026ldquo;database\u0026rdquo;, \u0026ldquo;router\u0026rdquo; and state machine all at once.\nConclusion One more benefit comes as a bonus with this architecture: it makes any LLM better at attention. A firing KS never sees the whole transcript, only a focused projection. The CQRS read model rendered for its narrow specialization and a single task. The context stays small not because it was truncated, but because it was projected. Nothing competes for attention that does not belong to the task. The \u0026ldquo;lost in the middle\u0026rdquo; problem is avoided.\nChat remains what it always should have been: an interface. Underneath, every message is a speech act journaled as an event. Every intent becomes an unresolved node that cannot be forgotten.\nDoes your agent still keep its world model in the transcript? Leave me a comment below!\nReferences J. L. Austin, How to Do Things with Words, Oxford: Clarendon Press, 1962. Allen Newell, \u0026ldquo;Some Problems of Basic Organization in Problem-Solving Programs\u0026rdquo;, in Yovits, Jacobi \u0026amp; Goldstein (eds.), Conference on Self-Organizing Systems, Spartan Books, 1962. John R. Searle, \u0026ldquo;A Taxonomy of Illocutionary Acts\u0026rdquo;, in K. Gunderson (ed.), Language, Mind, and Knowledge, University of Minnesota Press, 1975. Tim Finin, Richard Fritzson, Don McKay, Robin McEntire, \u0026ldquo;KQML as an Agent Communication Language\u0026rdquo;, Proceedings of the Third International Conference on Information and Knowledge Management (CIKM \u0026lsquo;94), 1994. Foundation for Intelligent Physical Agents, FIPA 97 Specification Part 2: Agent Communication Language, 1997. Qingyun Wu et al., \u0026ldquo;AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation\u0026rdquo;, arXiv:2308.08155, 2023. John R. Boyd, A Discourse on Winning and Losing, unpublished briefing, 1987. Nelson F. Liu et al., \u0026ldquo;Lost in the Middle: How Language Models Use Long Contexts\u0026rdquo;, Transactions of the Association for Computational Linguistics, 12, 2024. ","date":"2026-07-24T00:00:00+02:00","image":"/p/why-llm-chat-feels-like-an-anti-pattern/banner-b11_hu_ad8fcce89afb12da.png","permalink":"/p/why-llm-chat-feels-like-an-anti-pattern/","title":"Why LLM chat feels like an anti-pattern in 2026"},{"content":"There is a continuation of this article, Why LLM chat feels like an anti-pattern in 2026?, which shows the blackboard architecture in action.\nA blackboard architecture, modernized with event sourcing, content-addressed knowledge source (KS) firing and a few more enhancements.\nWhile working on improvements to various LLM-powered components and subsystems over the past year, I noticed I kept converging on the same architecture — one that, as I later found out, had been described and named four decades ago.\nThe blackboard model of problem solving is less known today not because it is so recent, but quite the contrary, because it was already summarized and canonized in 1986 by H. Penny Nii. The article \u0026ldquo;The Blackboard Model of Problem Solving and the Evolution of Blackboard Architectures\u0026rdquo; was published in two parts, the first part in AI Magazine, Volume 7, Issue 2 [1, 2].\nThat is right, 40 years ago, AI Magazine, long before we had GPUs pushing TOPS. And the metaphor itself is older still: the first \u0026ldquo;blackboard\u0026rdquo; in the AI literature appears in 1962, when Allen Newell imagined \u0026ldquo;a set of workers, all looking at the same blackboard: each is able to read everything that is on it, and to judge when he has something worthwhile to add to it\u0026rdquo; [3].\nThe model/architecture itself is relatively simple, based on a shared solution space (blackboard) and knowledge sources (humans, computers) participating in a solution. There is no classic \u0026ldquo;agentic loop\u0026rdquo; in this architecture. The set of knowledge sources is being evaluated by priority until none of them is firing. Each KS could plant new evidence on the board and the evaluation is potentially infinite.\nIt wasn\u0026rsquo;t a mere theoretical exercise back then. In fact, the article describes existing and functional systems such as HEARSAY-II (speech understanding) [5], HASP/SIAP (sonar signal interpretation) [6], CRYSALIS (protein structure inference), TRICERO (air-activity monitoring) and OPM (opportunistic errand planning), and summarizes theory and experience gathered in preceding decades.\nInterestingly, the article itself describes limitations of execution and control. These were partly theoretical and mostly practical limitations dictated by the available computers of that era. Footnote 7 of Part One [1] even treats the serialization of the model as a mere uniprocessor artifact and looks ahead to hundreds of processors.\nThe control decisions posted on the blackboard and made by the control KSs, scheduled like the domain KSs, were actually described a year earlier by Barbara Hayes-Roth (1985 article: \u0026ldquo;A Blackboard Architecture for Control\u0026rdquo;) [4]. McManus (1990) [8] later worked at the same architectural layer, contributing design and analysis metrics (output-overlap and I/O-connectivity between knowledge sources) for concurrent blackboard systems.\nMy Enhancements As I mentioned in the first sentence, the most prominent additions are event sourcing and content-addressed firing. These two choices work together to enable a third key enhancement: execution as a KS!\nThe 1980s systems kept execution in a separate action layer on purpose. They would be forced to re-run the side effects on every step (catastrophic to performance) or to build ad-hoc guards (the \u0026ldquo;did we already do this\u0026rdquo; kind).\nThe new event-sourced substrate [9] and N-tuple-based KS firing (a generalization of Rete\u0026rsquo;s refraction [7], with the content-addressed matching itself echoing Linda\u0026rsquo;s tuple spaces [10]) make the per-step evaluation cheap while guaranteeing the action fires once. The idea that evaluation cost should scale with the change rather than with the whole knowledge base has a long afterlife in incremental computation research and is very much alive today.\nIn today\u0026rsquo;s backend jargon, this is CQRS: the event journal is the write side and the blackboard is a projection — an ephemeral read model, never a mutable database. Replay merely reconstructs projections from journaled events; the executor cannot re-fire, because its past action is itself an event in the journal and the content-addressed gate sees that the question was already asked. The executor KS could be treated like every other KS and that is a huge win.\nA philosophical note on concurrency:\nlike SQLite, this is deliberately an engine, not a server. Where Nii\u0026rsquo;s footnote treated serialization as an artifact of uniprocessor hardware and looked ahead to hundreds of processors, I embrace it — deliberate serialization is exactly what buys replayability. The hard part of the problem is not throughput but the universe of possible solution trees; concurrency can stay somebody else\u0026rsquo;s problem.\nPractical Implications Beyond the benefits that originally motivated the design, I later realized the architecture also absorbs many of today\u0026rsquo;s constructs quite naturally.\nConsider the classic chat with an LLM: a KS observing the latest post from a human operator, invoking the LLM to respond. This sounds trivial until you realize:\nThe operator could respond to any previous node in the conversation — the KS will see an \u0026ldquo;unreplied-to\u0026rdquo; post and invoke the LLM to respond, effectively creating a branched conversation (and a chronological journal of events as a bonus). This capability became available only recently (and in various forms) with current coding harnesses and AI chat applications. The tree of conversation nodes provides an accurate path for each leaf and a cache-friendly prefix for any branched conversation. The introduction of multiple instances of the same KS (using different LLMs or varying system prompts) already creates a research team. Wrap any coding harness (such as Claude Code with the -p switch) into a KS and you have built an automated and insanely powerful development system. The unifying observation is that a blackboard system could become any current app: a chat, a loop, an orchestrator. It depends purely on the roster of enabled KSs and their vocabulary, and the realization of these patterns is elegant rather than merely possible.\nMain Benefits of a Blackboard Architecture Apart from the natural absorption of any common LLM app pattern used today, the blackboard architecture provides non-trivial benefits which are often \u0026ldquo;bolted on\u0026rdquo; within other systems.\nControl is decentralized and arbitration is coordinated. This strikes harder than it sounds: consider the case where a shell command is supposed to run a certain test suite of unit tests. The command could either succeed or fail (e.g. due to a wrong path), but the actual test execution must be assessed by a different KS. Did the test suite actually run? Did the test actually cover the source code properly? This leads to an important realization: There is no single result of an action — it all depends on \u0026ldquo;who is watching\u0026rdquo;, and thus any other system with a single task outcome will be severely limited in properly and accurately representing the state of the world. Event sourcing provides temporal consistency, replayability, audit trails and natural versioning all at the same time. Blackboard as a projection avoids uncontrolled state mutation and consistency issues. The projection also doesn\u0026rsquo;t have to be a single monolithic board: the same journal can materialize into specialized read models per KS — for example straight into a search index for a RAG-supporting KS. Content-based KS firing maps cleanly to reactive/subscription patterns and thus to many common products used today (Kafka topics, pub/sub or even rule engines) — a decoupling already fully formed in 1985 in Linda\u0026rsquo;s tuple spaces [10]. Heterogeneous knowledge sources are natural: Prominently LLM-based KSs, either in a pure form (KS code invoking an LLM directly) or by embedding another product (such as Claude Code in non-interactive mode). The context rendering is just another projection which could be tailored for every kind of KS separately if desired. A human operator as a KS becomes \u0026ldquo;human in the loop\u0026rdquo; in a powerful way. Their view of the \u0026ldquo;world\u0026rdquo; could be implemented as a UI; the underlying substrate is still the single source of truth. A whole gradient of symbolic reasoners, simulators, sandbox environments, build pipelines, connectors to databases (RAG-supporting KSs), etc. Recursive composability: The blackboard system could be using knowledge sources which are themselves embedded blackboard systems with a specific (focused) set of KSs and specific vocabulary. Limitations There is a potential trap in the decentralized control inherently embedded in a set of knowledge sources. This was historically also the reason why blackboard systems with large rule sets became unmaintainable. This should be avoided and proper focus and decomposition are advised. Instead of building a set of KSs which could do anything, it is better to focus on a set of KSs addressing a particular task exceptionally well. No real-time systems with any kind of latency requirements. A blackboard system is better thought of as a team of experts, built to solve difficult problems without any time expectations. Like with Turing completeness, many architectures could be simulated as a blackboard system, but that does not mean it is the optimal or idiomatic fit. Conclusion Perhaps the best way to summarize the whole design: this is not just a blackboard. It is a continuously evolving, historically accurate database of reasoning.\nHave you converged on a similar architecture in your own work, or do you disagree with my design choices? Please, leave me a comment below!\nReferences H. Penny Nii, \u0026ldquo;The Blackboard Model of Problem Solving and the Evolution of Blackboard Architectures\u0026rdquo; (Part One), AI Magazine, 7(2), 1986. H. Penny Nii, \u0026ldquo;Blackboard Application Systems and a Knowledge Engineering Perspective\u0026rdquo; (Part Two), AI Magazine, 7(3), 1986. Allen Newell, \u0026ldquo;Some Problems of Basic Organization in Problem-Solving Programs\u0026rdquo;, in Yovits, Jacobi \u0026amp; Goldstein (eds.), Conference on Self-Organizing Systems, Spartan Books, 1962. Barbara Hayes-Roth, \u0026ldquo;A Blackboard Architecture for Control\u0026rdquo;, Artificial Intelligence, 26(3), 1985. Lee D. Erman, Frederick Hayes-Roth, Victor R. Lesser, D. Raj Reddy, \u0026ldquo;The HEARSAY-II Speech-Understanding System: Integrating Knowledge to Resolve Uncertainty\u0026rdquo;, ACM Computing Surveys, 12(2), 1980. H. Penny Nii, Edward A. Feigenbaum, John J. Anton, A. J. Rockmore, \u0026ldquo;Signal-to-Symbol Transformation: HASP/SIAP Case Study\u0026rdquo;, AI Magazine, 3(2), 1982. Charles L. Forgy, \u0026ldquo;Rete: A Fast Algorithm for the Many Pattern/Many Object Pattern Match Problem\u0026rdquo;, Artificial Intelligence, 19(1), 1982. John W. McManus, \u0026ldquo;Design and Analysis Tools for Concurrent Blackboard Systems\u0026rdquo;, 9th IEEE/AIAA Digital Avionics Systems Conference, NASA Langley Research Center, 1990. Jay Kreps, \u0026ldquo;Questioning the Lambda Architecture\u0026rdquo;, O\u0026rsquo;Reilly Radar, 2014. David Gelernter, \u0026ldquo;Generative Communication in Linda\u0026rdquo;, ACM Transactions on Programming Languages and Systems, 7(1), 1985. ","date":"2026-07-06T00:00:00+02:00","image":"/p/what-is-better-than-your-ai-loop/cover_hu_3c3546198588fe5c.png","permalink":"/p/what-is-better-than-your-ai-loop/","title":"What is better than your AI loop?"},{"content":"It\u0026rsquo;s been over a year since I published my view on the state of AI/LLMs and how it could be significantly improved in my post: \u0026ldquo;Failure Is Not An Option For AI\u0026rdquo;. While scrolling through the rStar2-Agent technical report, I couldn\u0026rsquo;t help but mumble to myself: \u0026ldquo;I told you so!\u0026rdquo;.\nConsider how quickly the model improved during training, especially given that we are comparing a 14B parameter model to 671B R1.\nThis doesn\u0026rsquo;t just show how vast the universe of optimization is; it demonstrates the fundamental principle:\n\u0026#x27a1;\u0026#xfe0f; The cost of inaction often far exceeds the cost of a reversible mistake.\nChain of Action When comparing the traditional Chain-of-Thought (CoT) approach to this new \u0026ldquo;Chain of Action\u0026rdquo;, the most striking difference is how early the environmental feedback provides value to the training process. To appreciate what was achieved, let\u0026rsquo;s look at technical side. Using this approach effectively during training required a high-throughput execution environment. The one used for this project was capable of handling 45,000 concurrent tool calls, returning feedback in just 0.3s on average.\nCompounding is a fundamental principle of investing. Becoming the richest person in the world is possible by investing small sums and compounding them well enough and long enough. But this principle works inefficiently when applied to chain of thought. Compounding subtle errors early in a reasoning process leads to a long, inefficient, and ultimately incorrect reasoning trajectory.\nAnyone who sat through lengthy corporate meetings or a drawn-out group decision-making process understands this intuitively. The first practical attempt provides more knowledge than a month of theoretical planning. Feedback, even when negative, is crucial and must be delivered swiftly. I cannot stress this enough.\nLearning from Negative Feedback The emergent ability of the model trained with this novel approach is particularly unique: The model learned to react productively to negative feedback. Researchers observed the model using specific \u0026ldquo;forking\u0026rdquo; and \u0026ldquo;reflection\u0026rdquo; tokens. It was effectively talking to itself - course-correcting, pausing to analyze an error, exploring alternative approaches.\nThis suggests a universal success formula, for humans and AI alike:\n\u0026#x27a1;\u0026#xfe0f; Form a hypothesis, take action, observe feedback, and repeat.\nThe best part of this story is that the rStar2-Agent codebase has been released under an MIT license on GitHub.\nReferences Microsoft Research, \u0026ldquo;rStar2-Agent: Agentic Reasoning Technical Report\u0026rdquo;, arXiv:2508.20722, 2025. rStar2-Agent repository, GitHub, MIT licence. ","date":"2025-09-08T00:00:00+02:00","image":"/p/why-success-favors-action-and-ai/cover-image01_hu_149ccc0fbda131a6.png","permalink":"/p/why-success-favors-action-and-ai/","title":"Why Success Favors Action and How This Relates to AI"},{"content":"Numerous copy routine implementations are readily available in .NET. If I were to simply list them alongside a few benchmark numbers and charts, it wouldn\u0026rsquo;t make for a very interesting article.\n\u0026#x26a0;\u0026#xfe0f; What if I told you upfront that none of these routines is designed to be the absolute fastest?\nIf you\u0026rsquo;re interested in a basic comparison, I recommend checking out this article here on dev.to: What is the best way to copy an array? or a more detailed, older comparison: High performance memcpy gotchas in C#.\nHere, I\u0026rsquo;ll outline a list of options, but this is far from the whole story:\nA simple for loop (hint: foreach is usually a bit faster) Array.Copy Span.CopyTo Buffer.BlockCopy Buffer.MemoryCopy Marshal.Copy Unsafe.CopyBlock Imported memcpy If you\u0026rsquo;re currently struggling with slow array or memory copy operations, try one of the functions on this list.\nHowever, there are some elephants in the room, and I plan to uncover a few of them.\nElephant No.1 - Cache Pollution What You might read in many places is that framework functions are already highly optimized. This is true, but they are not necessarily optimized for the highest speed possible.\nBuilt-in functions provided by .NET framework are optimized in various ways. One significant consideration is preventing cache pollution. Wait - what? Yes, x86 CPUs achieve their speed primarily thanks to their cache. If the cache is disabled or not utilized properly, code execution speed drops dramatically.\nCache Pollution Explained Cache pollution when copying large data blocks: This occurs when frequently used data gets replaced in the CPU cache by the data being copied. Chances are that once You are done copying, You won\u0026rsquo;t touch the same data ever again, so it makes placing them in a CPU cache unnecessary. For example, imagine a network stack - once data is sent, it\u0026rsquo;s unlikely to be touched again. Similarly, when loading textures with the CPU for GPU usage, caching may be unnecessary.\nStandard .NET functions mitigate this issue when copying larger blocks (usually sizes above 1MB) by using non-temporal access, bypassing the cache.\nThis might not suit your use case. For instance, if you are waiting for the copy to complete before doing anything else - some cache pollution would be an acceptable tradeoff. Especially when the data is likely to be cached anyway, as in simulations or CPU rendering.\n\u0026#x25b6;\u0026#xfe0f; How can we observe this? And what can we do about it?\nChart above compares various buffer sizes (ranging from 1MB to 100MB) and copy methods. The leftmost data points for 1MB block shows great performance for Buffer.MemoryCopy and Unsafe.CopyBlock, two best methods available in .NET for memory copying. However performance falls sharply past 1MB.\nTo illustrate the real HW limitation, please notice the comparison between AVX based copy routine using 256bit (orange) and 512bit (red) vectors and the difference between normal loads and non temporal 512 bit variant (light blue). These tests were run on a Ryzen 9950x, with 48kB L1 data cache, 1024kB of L2 and 32MB of L3 available to any single CPU core. Total cache sizes are 1280 KB L1 Cache (16x 32kB + 16x 48kB), L2 Cache 16 MB (16x 1024kB) and L3 Cache 64 MB (2x32MB).\nUsing cached AVX loads, high copy speed is sustained even for 8MB (and larger) blocks. However, when copying large blocks relative to cache size (e.g.: 100MB), built-in functions regained their advantage.\nSame data, other views — linear and logarithmic axes:\nElephant No.2 - Overhead My first benchmark shows that splitting larger buffers into smaller chunks can improve performance at the expense of increased CPU cache utilization.\nThere are other factors causing slowdowns, for example:\nManaged memory introduces framework checks for each access. Unaligned memory access could decrease the CPU cache efficiency. And finally, with unmanaged, 64 byte aligned memory, there is still a final question: \u0026#x25b6;\u0026#xfe0f; What is the optimal block size?\nThis depends on the actual CPU, but we are trying to strike a balance between call overhead and efficient CPU resource utilization.\nThe chart above shows the throughput for transferring a 32MB buffer using various block sizes and methods. The key takeaway is Buffer.MemoryCopy (blue) and Unsafe.CopyBlock (yellow) perform best with block sizes between 8kB to 1MB. Notably, these methods outperform themselves (orange), compared with single call for the whole 32MB buffer.\nMethods represented by horizontal lines do not use variable block sizes. It is worth mentioning that AVX variants are always loading 8 vectors (of 256 or 512 bits) before storing them, thus effectively working with 256 an 512 BYTE blocks regardless of the total buffer size.\nThe 32 MB transfer on a linear and on a logarithmic block-size axis:\nElephant No.3 - Multi Threading So far, we have tested only a single threaded performance, as memory speed and cache capacity are often the limiting factor. However, we are not utilizing all CPU resources, which might cost us some performance!\n\u0026#x25b6;\u0026#xfe0f; Can we combine the previous techniques with multiple threads?\nModern CPUs, even desktop ones, are becoming increasingly heterogeneous. By splitting the workload across multiple threads, we might take advantage of more CPU resources:\nThe chart above shows the chunked, multi-threaded approach. For reference, the test system uses dual channel DDR5 6000MT/s memory, with theoretical peak performance about half of the total throughput (50% of 90GB/s).\nWhat we observe here is a synergistic effect between smaller blocks and multiple threads!\nLarger versions of the multi-threaded chart and of the per-method comparison:\nFinal Thoughts Standard framework functions are well-optimized, safe and sufficient for most use cases. If extreme performance is required, there are additional techniques and tradeoffs available. Example with a single call to Buffer.MemoryCopy on an 8MB buffer reaching ~30GB/s speed, offers significant gains. It is possible to reach 220GB/s with multiple transfers of 128kB blocks and using multiple threads on a same CPU. This is more than 7x improvement.\nReferences j0nimost, \u0026ldquo;What is the best way to copy an array?\u0026rdquo;, dev.to. \u0026ldquo;High performance memcpy gotchas in C#\u0026rdquo;, Code4k blog, 2010. ","date":"2024-12-20T00:00:00+01:00","image":"/p/fast-memory-copying-in-csharp/cover_hu_3f8d3b6539b2c16a.png","permalink":"/p/fast-memory-copying-in-csharp/","title":"Fast Memory Copying in C#/.NET (Cache, AVX, Threads, Unsafe, Alternatives)"},{"content":"There\u0026rsquo;s something uniquely captivating about revisiting the passions of our past. I still remember the countless hours spent tinkering with code on my old 386 and 486 PCs, mesmerized by the magic of pixels coming to life on the screen. The simplicity of those days, where an array of pixels and a spark of imagination were all you needed, still holds a special place in my heart. Inspired by that nostalgia, I decided to breathe new life into one of my early projects - a 2D firework simulation - using modern tools and techniques. I am using Mode-13HX as a basis today, for more details please check my previous part: Old-School Graphics in C# / .Net 8, Part 1.\nHistory This was originally written in (or before) a year 2000 using Pascal (programming language) and compiled using Turbo Pascal. It used video mode 13h under MS-DOS. Originally 320x200 pixels @ 8bit palette / 256 colors. I have ported this project multiple times in last 20+ years, since it is a small and fun project to play with.\nI am running it at 4k today, thanks to modern C# and Mode-13HX. It should work on any platform and architecture where .NET 8 and C# is supported. I added special code path for AVX512 (fading) and SSE (color mixing) to speed things up - If You are interested in these, please keep reading, scroll right to code samples or check the GitHub repo right away.\nFun with Particles The core of the simulation revolves around a 2D particle system that emulates fireworks bursting in the night sky. Each firework is composed of multiple particles that are generated, animated, and eventually discarded.\nInitial Launch: Particles are spawned with initial positions, velocities and countdown timer that simulate the upward launch of a firework rocket. Explosion Stages: Each particle has a timer determining its lifespan, ensuring that old particles are removed to make way for new ones. When time is out, particles could burst into additional particles (or disappear), creating multiple stages of explosion for a more dynamic display. Gravity and Air: Somewhat realistic motion is achieved by applying gravitational acceleration and air resistance, affecting each particle\u0026rsquo;s trajectory over time. Original simulation expected constant framerate and particle deceleration was expressed as a relative speed loss over constant time period. For example 85% of speed after 1/60s. In order to account for the variable (or unknown constant) frame rate, I recalculated this into a look-up table (velocity and a position decay factors) for all possible frame offsets with a microsecond resolution. Maximum frame offset is capped, so we don\u0026rsquo;t run out of pre-calculated table. To create a vibrant and immersive visual experience, the simulation uses additive color blending.\nAdditive Blending: When particles overlap, their colors are added together, increasing brightness and creating a glow effect. Flare Rendering: Each particle is drawn not just as a single point but with a small flare (1+4 pixels) to enhance the luminosity and visual appeal. Color Sets: Multiple predefined color sets (e.g.: standard RGB, pastel tones) are used to diversify the fireworks\u0026rsquo; appearance. A critical aspect of the simulation is the gradual fading of framebuffer to simulate the dissipating light of fireworks.\nFrame Buffer Fading: The entire frame buffer undergoes a fading process where pixel brightness decreases over time. This is actually quite inefficient approach because we need to swap at least 2 back buffers. This was not the case in DOS days, but this time the particle drawing and fading is done using a separate (+1 additional) buffer. Performance Optimizations with Vectorization To maintain high performance, especially when rendering thousands of particles, the simulation employs vectorization techniques using SIMD (Single Instruction, Multiple Data) instructions available in modern CPUs. While it may look obvious to use vectorization to accelerate the particle movement itself, I opted not to do so. In terms of CPU usage, the most expensive parts are color blending, fading and page flipping (copying of a framebuffer into texture buffer). It might be an interesting exercise to rewrite this whole thing into OpenCL though. But, Hey! Maybe here is the place for a link to Your repository!\nSSE Optimization for Color Mixing. When blending colors for overlapping particles, SSE instructions add multiple color components (red, green, blue - 8bits each) in parallel. One funny (or sad) fact: only 24bits out of 128bit vector are used effectively, but it is still worth the performance increase!\nprivate static void MixColors(ref uint target, uint color) { if (Ssse3.IsSupported) { Vector128\u0026lt;uint\u0026gt; targetVector = Vector128.CreateScalar(target); Vector128\u0026lt;uint\u0026gt; colorVector = Vector128.CreateScalar(color); Vector128\u0026lt;byte\u0026gt; targetBytes = targetVector.AsByte(); Vector128\u0026lt;byte\u0026gt; colorBytes = colorVector.AsByte(); Vector128\u0026lt;byte\u0026gt; resultBytes = Sse2.AddSaturate(targetBytes, colorBytes); target = resultBytes.AsUInt32().ToScalar(); } else { // Fallback to non-SIMD code } } Frame Buffer Fading: The fading effect subtracts a small value from each color channel across all pixels. AVX-512 accelerates this by handling multiple (16) pixels simultaneously. This reduces the CPU load and ensures that the fading effect doesn\u0026rsquo;t become a performance bottleneck, even at high resolutions.\nprivate const int VectorSize = 64; private void FadeScenePixels(uint[] pixels, double timePassed) { if (Avx512BW.IsSupported) { int totalBytes = pixels.Length * sizeof(uint); int fadeValue = CalculateFadeValue(timePassed); Vector512\u0026lt;byte\u0026gt; fadeVector = Vector512.Create((byte)fadeValue); unsafe { fixed (uint* pScenePixel = pixels) { byte* p = (byte*)pScenePixel; for (int i = 0; i + VectorSize \u0026lt;= totalBytes; i += VectorSize) { Vector512\u0026lt;byte\u0026gt; pixelVector = Avx512BW.LoadVector512(p + i); pixelVector = Avx512BW.SubtractSaturate(pixelVector, fadeVector); Avx512BW.Store(p + i, pixelVector); } } } } else { // Fallback to non-SIMD code } } While it would be possible to continue with more optimizations and enhancements, I believe this might be a good stopping point for now. I like to keep fun projects small and simple, so they keep being funny.\nSource code: Fireworks 2D GitHub repository.\nReferences fireworks2D, the 2D firework simulation (C#, .NET 8, SSE/AVX-512 paths), GitHub. mode-13hx, the Mode 13HX template project it builds on, GitHub. Pascal (programming language) and Turbo Pascal, Wikipedia. ","date":"2024-11-20T00:00:00+01:00","image":"/p/old-school-graphics-part-2-fireworks-avx-sse/banner-fw2d_hu_611a8b301587660.png","permalink":"/p/old-school-graphics-part-2-fireworks-avx-sse/","title":"Old-School Graphics in C# / .Net 8, Part 2: Fireworks and Advanced Vector Extensions (AVX, SSE)"},{"content":"I began my journey with computers and programming during the era of 386 PCs, a time when DOS games ruled and sparked endless curiosity. I remember playing those games and wondering how they were created. It wasn’t long before I discovered Mode 13h - a graphics mode that perfectly fit within a 64kB page in 16-bit code, thanks to its 8 bits per pixel, 256-color palette, and 320×200 resolution.\nMany years and APIs have passed since, but every now and then, I feel nostalgic for that simpler time. Back then, all we had was an array of pixels and an idea - and that was all we needed. I recall the breathtaking graphical demos that showcased incredible creativity, often categorized by their size. Some were only a few kilobytes but managed to display complex scenes with textures, geometry, animations, and even synchronized music.\nThat era had a certain magic. Many of the games of that time are legendary, and are still being ported to modern platforms and APIs today. Revisiting that period feels like embarking on an archaeological adventure - uncovering hidden treasures in the form of old code, design principles, and mathematical tricks. While reliving those times, I often reflect on the stark contrast with today: CPUs are now far more powerful than the high-end GPUs of the early 2000s, offering astounding possibilities. Yet, with all this power, we’ve lost much of the simplicity in the creative process.\nTo bring back the spirit of Mode 13h, I decided to create a template C# project which I named Mode 13HX. The goal was to capture the simplicity of the past while embracing modern practicality. There’s no need to limit ourselves to 256 colors or 320×200 resolution anymore. Instead, I focused on supporting more platforms (x86, ARM) and operating systems (Windows, Linux, macOS). This project provides a simple yet powerful entry point, paying homage to the roots of graphics programming while adapting to today’s possibilities.\nHistory Mode 13h is a graphics video mode introduced with IBM\u0026rsquo;s VGA (Video Graphics Array) standard in 1987. It provides a resolution of 320×200 pixels with 256 colors, marking a significant advancement in PC graphics capabilities at the time. Due to its straightforward programming model and direct access to video memory, mode 13h became popular among game developers in the late 1980s and early 1990s. It played a crucial role in the evolution of PC gaming, allowing for more detailed and colorful graphics despite its relatively low resolution.\nTechnical Details:\nMode Number: BIOS video mode 0x13 Resolution: 320×200 pixels with a 16:10 aspect ratio, approximating 4:3 on CRT monitors Color Depth: 8 bits per pixel (256 colors) from a palette of 262,144 (18-bit RGB) Video Memory Location: Linear framebuffer starting at segment 0xA000 Memory Usage: 64 KB of video memory Access Method: Direct memory mapping without bank switching Mode 13HX When re-imagining this mode, it’s no longer necessary to constrain ourselves with a limited palette or small resolution - these constraints can be easily emulated if needed. In my opinion, the most compelling aspect is the ability to directly access, modify, and manage a continuous block of pixels in memory before displaying it.\n+-----------------+-----------------+-----------------+ | | | | | Frame 0 | Frame 1 | Frame 2 | \u0026lt;-- CPU writes pixels | | | | IRasterizer.Render() +--------+--------+-----------------+-----------------+ | V +--------+--------+ +-------------------------------+ | OpenGL Texture | | Vertex buffer | | (updated by FB) | | (with texture coordinates) | +--------+--------+ +--------------+----------------+ | | V V +---------------------------------------------------------------+ | OpenGL Rendering by GPU | | (Draw Calls) | +---------------------------------------------------------------+ | V +-------------+ | Screen | +-------------+ I chose the great OpenTK toolkit and its OpenGL bindings. The idea here is simple: setup a textured polygon (two triangles) and make sure it covers the viewport exactly while updating its texture for each frame. Technically we have the power of GPU on our side as well. This is especially useful when we need to scale the image up (or down) to match the screen resolution. The texture data we update is essentially a frame from our framebuffer!\nTechnical Details \u0026amp; Usage:\nBase C# project: Provides examples, IRasterizer interface and direct pixel access. Built with OpenTK targeting .NET 8, tested to work on Windows and Linux. GitHub repository: mode-13hx Resolutions: default is 1920×1080 (Full HD), customizable via -w, -h options, works with any resolution. Tested with 3840×2160 (4K). Modes: full screen (-f) and windowed (default). V-Sync Support: Enable with -v to synchronize rendering with display refresh rate. Color Depth: 24-bit RGB (True Color), supporting 16.7 million colors. Each pixel uses 3 bytes (RGB), eliminating the need for palettes. Framebuffer: Linear framebuffer for direct pixel manipulation. Double Buffering: Default method (parameter -l 1 is default). Triple Buffering: Optional (-l 2 or more) for higher fps. Rendering: Direct pixel manipulation by the CPU in the framebuffer. Frame is rendered onto the screen using OpenGL as a texture. Command-Line Interface: dotnet mode13hx.dll [command] [options] Input Handling: Simplified input via OpenTK for keyboard and mouse support. Performance: Supports high frame rates: 60, 120, 144 FPS with V-Sync; uncapped without V-Sync. Repository: mode-13hx\nReferences mode-13hx, the Mode 13HX template project (C#, .NET 8), GitHub. OpenTK, .NET bindings for OpenGL with windowing and input. ","date":"2024-11-13T00:00:00+01:00","image":"/p/old-school-graphics-part-1-mode-13hx/cover_hu_def9f15da43960a5.png","permalink":"/p/old-school-graphics-part-1-mode-13hx/","title":"Old-School Graphics in C# / .Net 8, Part 1: Teaching an Old Dog New Tricks (Introducing Mode 13hx)"},{"content":"Zen 5 landed with a bit of controversy, but anyone Who was paying any attention to the real results wasn\u0026rsquo;t disappointed. Benchmarks done using more mature OSes (like Linux and Win10) and professional software shown its strengths. But hey, why should we trust just any 3rd party data if we can make a test by ourselves? I used Ryzen 9950x today and a bit of code to get things going.\nData Processing and Core-to-Core Latency I am often processing large piles of data. While this is normally executed on different HW than my personal PC, it is always nice to have the computing power locally.\nWhat is counterintuitive in a case of data processing: more CPU cores may not bring better performance. In a CPU heavy tasks (like rendering) this is not apparent, but when the CPU work is combined with synchronization and data transfers the core to core latency, data locality, OS scheduler and other thigs are coming into play.\nThe Producer–Consumer Test Setup I promised a practical example, so lets examine this old problem of producer and consumer. We will dive into details soon (or jump right into the code here), but now lets focus on the main code part. I am using C# with .NET 8.0 and Win10.\nIntPtr mask1 = new(1 \u0026lt;\u0026lt; core1); IntPtr mask2 = new(1 \u0026lt;\u0026lt; core2); SemaphoreSlim semaphoreProducer = new(queueSize); SemaphoreSlim semaphoreConsumer = new(0); string[] message = new string[queueSize]; int reader = -1; int writer = -1; In the code above we define 2 affinity masks which are binary vectors of ones and zeroes hidden inside single int type. We allow only one core in each mask, values for core1 and core2 are within the range 0..15 (I disabled SMT for this test, otherwise we would have to deal with virtual cores as well). I use two semaphores for tight control of producer and consumer threads. The producer could produce only up to queueSize messages and consumer could consume only as many messages as are available without blocking. Array message is our queue and reader/writer index is used to address next message (or empty spot) in our queue.\nProducer Thread producerThread = new(() =\u0026gt; { ThreadGuard.GetInstance(mask1).Guard(); for (int i = 1; i \u0026lt;= messagesToProcess; i++) { semaphoreProducer.Wait(); int index = Interlocked.Increment(ref writer); message[index % queueSize] = i.ToString(); semaphoreConsumer.Release(); } }); Here comes the producer thread. First thing we do is the affinity setting. ThreadGuard implementation is a technical detail for which You have to scroll a bit further, it ensures that the thread it is called from will get executed only on the core(s) we defined by a mask. We produce desired number of messages in a for loop. Each message needs it own place in a queue so we do semaphoreProducer.Wait() as a first thing. Initial capacity of this semaphore is queueSize, so there is plenty of place from the start. We are looping through spots in a queue thanks to modulo operation, thus index % queueSize. The call semaphoreConsumer.Release() is actually telling the consumer about the new message we just prepared.\nConsumer Thread consumerThread = new(() =\u0026gt; { ThreadGuard.GetInstance(mask2).Guard(); int i = 0; do { semaphoreConsumer.Wait(); int index = Interlocked.Increment(ref reader); i = int.Parse(message[index % queueSize]); semaphoreProducer.Release(); } while (i \u0026lt; messagesToProcess); }); Consumer part looks like a mirror of the producer, we wait in semaphoreConsumer.Wait() until there is some message to consume, once we are done the semaphoreProducer.Release() call signals the producer that the place in queue is now free for a new message. Whole party is going on until we produce and consume desired amount of messages.\nThreadGuard Complete ThreadGuard code looks like this:\nusing System; using System.Collections.Concurrent; using System.Runtime.InteropServices; public class ThreadGuard { // Import SetThreadAffinityMask from kernel32.dll [DllImport(\u0026#34;kernel32.dll\u0026#34;)] private static extern IntPtr SetThreadAffinityMask(IntPtr hThread, IntPtr dwThreadAffinityMask); // Import GetCurrentThread from kernel32.dll [DllImport(\u0026#34;kernel32.dll\u0026#34;)] private static extern IntPtr GetCurrentThread(); private static readonly ConcurrentDictionary\u0026lt;IntPtr, ThreadGuard\u0026gt; GuardPool = new(); private readonly IntPtr mask; private ThreadGuard(IntPtr mask) { this.mask = mask; } public static ThreadGuard GetInstance(IntPtr affinityMask) // factory method { return GuardPool.GetOrAdd(affinityMask, _ =\u0026gt; new ThreadGuard(affinityMask)); } public void Guard() { IntPtr currentThreadHandle = GetCurrentThread(); // Get the handle of the current thread IntPtr result = SetThreadAffinityMask(currentThreadHandle, mask); // Set the thread affinity mask if (result == IntPtr.Zero) { throw new InvalidOperationException(\u0026#34;Failed to set thread affinity mask.\u0026#34;); } } } Results Now the best part - let\u0026rsquo;s run this for all pairs of cores and various queue sizes. All measured numbers are millions of transferred messages per second. Row numbers on a left are \u0026ldquo;producer\u0026rdquo; cores and column numbers on top are \u0026ldquo;consumer\u0026rdquo; cores.\nQueue Length 1 Table 1 — 250k messages per pair, queue length 1. Millions of messages per second; rows = producer core, columns = consumer core.\nCoreA \\ CoreB -\u0026gt; 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 0.21 3.49 4.05 3.85 3.41 3.74 3.35 3.40 0.82 0.81 0.78 0.80 0.77 0.74 0.84 0.73 1 3.89 0.19 3.32 3.76 3.63 3.54 3.64 4.03 0.84 0.90 0.98 0.87 0.88 0.68 0.90 0.70 2 3.61 2.52 0.19 3.34 4.08 3.58 2.91 4.01 0.90 0.82 0.93 1.04 0.81 1.05 0.83 1.08 3 4.11 2.43 4.04 0.19 3.76 3.54 3.99 3.32 0.89 0.88 0.93 0.88 0.99 0.91 0.99 0.80 4 3.08 3.87 3.56 3.55 0.19 3.49 4.02 3.50 0.86 0.95 0.91 0.92 1.09 0.84 0.87 0.88 5 3.56 3.48 3.30 3.82 3.48 0.19 3.69 4.11 0.88 1.08 0.99 0.95 0.93 0.90 0.81 0.88 6 3.63 3.47 3.45 3.63 3.50 3.46 0.19 4.26 0.94 0.97 0.95 1.09 0.89 1.16 0.95 0.88 7 3.05 3.08 4.71 3.26 3.53 3.36 3.49 0.19 0.95 0.91 0.94 0.89 0.93 0.76 0.85 1.01 8 0.82 1.16 0.91 0.79 0.87 0.83 1.00 0.79 0.18 3.75 4.00 3.41 3.17 2.99 3.48 3.60 9 0.76 0.75 0.82 0.97 0.90 1.04 0.85 0.81 3.89 0.18 3.66 3.63 3.01 3.55 3.17 3.90 10 0.80 1.02 0.81 0.82 0.83 0.76 0.77 0.74 3.70 4.45 0.18 3.23 4.77 3.49 2.78 4.11 11 0.78 0.87 0.83 0.86 0.89 0.81 0.80 1.00 3.37 3.60 4.04 0.18 4.40 3.58 3.57 2.98 12 0.75 0.79 1.01 0.83 0.69 0.82 0.76 0.85 3.25 3.67 3.18 4.23 0.20 3.88 3.78 3.36 13 0.89 0.93 0.98 0.84 1.04 0.88 0.90 0.88 3.46 3.54 3.83 3.81 3.39 0.20 3.12 3.81 14 0.78 0.83 0.77 0.89 0.84 0.74 0.90 0.78 3.56 3.08 3.39 3.43 3.33 3.47 0.20 4.06 All five tables as colour-coded heat maps (green = fastest, yellow = same CCD, orange = cross-CCD), the way they were originally published — the plain tables follow below:\nStarting with queue length of 1 we are not letting any thread to produce more than 1 message without blocking. This looks interesting. Even worse than talking to another CCD is the single core trying to execute two things at once. I am surely not alone who hates to do context switching while working on multiple tasks. Solution to this (as visible later) is to at least do more work on a single task before switching to another one. For those who still wonder why is this so bad - we are really not allowing any thread to do any useful work until the other thread gets its fair share of CPU time. Measured throughput of messages is at the same time also measuring number of context switches between consumer and producer threads by the OS when only single CPU core is used. When we shift our attention to the single vs. cross CCD communication using 2 cores, it comes about 5x slower.\nQueue Length 10 and 100 Table 2 — 250k messages per pair, queue length 10.\nCoreA \\ CoreB -\u0026gt; 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 1.63 5.47 5.45 5.25 5.62 5.31 5.37 4.73 1.19 1.23 1.29 1.23 1.29 1.26 1.12 1.35 1 5.97 1.76 5.00 6.42 4.57 5.99 5.32 6.18 1.62 1.29 1.24 1.04 1.42 1.00 1.15 1.09 2 5.97 6.07 1.75 5.89 7.34 5.82 5.41 5.98 1.35 1.30 1.15 1.44 1.56 1.32 1.46 1.16 3 5.72 4.99 6.05 1.76 5.52 6.01 6.25 5.90 1.33 1.29 1.24 1.44 1.23 1.15 1.31 1.29 4 5.82 5.56 6.04 5.28 1.76 5.53 6.13 5.92 1.25 1.45 1.32 1.60 1.30 1.64 1.05 1.39 5 5.01 5.56 6.09 6.85 6.19 1.73 5.71 5.70 1.42 1.30 1.20 1.36 1.34 1.16 1.32 1.35 6 5.61 5.26 5.56 6.00 5.54 5.84 1.74 5.55 1.35 1.64 1.53 1.16 1.16 1.16 1.32 1.36 7 4.99 5.62 5.82 5.87 5.51 5.87 5.84 1.73 1.23 1.36 1.45 1.07 1.42 1.39 1.34 1.21 8 1.01 1.23 1.27 1.24 1.07 1.34 1.36 1.30 1.69 5.41 4.85 6.10 5.66 6.14 5.35 5.01 9 1.25 1.24 1.09 1.25 1.19 1.23 1.27 1.12 5.66 1.70 6.69 5.84 4.96 5.50 5.11 5.04 10 1.29 1.23 1.14 1.28 1.39 1.09 1.13 1.08 5.66 5.01 1.69 5.99 6.38 5.90 7.00 5.33 11 0.92 1.11 1.30 1.44 1.31 1.10 1.34 1.26 5.39 5.42 5.74 1.70 5.57 5.66 4.86 5.39 12 1.29 1.61 1.19 1.09 1.15 1.32 1.00 1.41 5.96 5.95 5.95 5.32 1.69 5.60 5.80 5.36 13 1.09 1.04 1.49 1.30 1.47 1.29 1.24 1.24 4.52 5.23 5.37 6.16 5.52 1.69 5.87 6.41 14 1.08 1.02 1.13 1.23 1.26 1.19 1.05 1.29 5.96 4.83 5.16 5.38 5.70 5.80 1.66 5.62 Table 3 — 250k messages per pair, queue length 100. CoreA \\ CoreB -\u0026gt; 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 8.27 5.54 5.21 5.39 5.14 5.48 4.93 4.82 1.33 1.30 1.34 1.17 1.32 1.36 1.30 1.40 1 6.19 10.05 6.16 5.31 6.20 5.50 5.72 4.76 1.41 1.35 1.40 1.33 1.34 1.41 1.44 1.41 2 4.63 4.59 10.00 6.51 6.23 5.43 6.11 5.20 1.58 1.51 1.51 1.62 1.46 1.33 1.25 1.32 3 5.75 5.84 6.23 9.92 5.57 6.05 6.02 5.41 1.42 1.18 1.45 1.37 1.12 1.43 1.26 1.15 4 5.76 5.91 6.13 5.42 10.00 6.43 5.65 5.64 1.57 1.20 1.23 1.33 1.38 1.30 1.15 1.41 5 5.33 6.10 5.58 6.27 6.29 9.99 5.31 5.86 1.19 1.31 1.52 1.38 1.37 2.04 1.20 1.33 6 5.73 5.30 5.38 5.61 5.93 6.12 9.90 6.00 1.76 1.61 1.18 1.19 1.32 1.51 1.21 1.35 7 5.39 6.25 5.69 5.73 5.73 5.98 5.81 9.81 1.29 1.67 1.46 1.46 1.09 1.23 1.32 1.35 8 1.33 1.13 1.31 1.43 1.19 1.46 1.32 1.37 9.59 5.99 6.18 6.06 5.88 5.66 5.59 5.56 9 1.15 1.34 1.41 1.41 1.36 1.35 1.06 1.21 6.02 9.64 5.41 5.24 5.12 5.05 5.12 5.50 10 1.31 1.23 1.00 1.13 1.42 1.33 1.32 1.15 5.47 6.33 9.73 6.00 5.63 5.28 5.77 5.48 11 1.17 1.29 1.57 1.25 1.28 1.19 1.60 1.10 5.33 6.29 5.71 9.60 5.83 6.21 5.31 5.37 12 1.15 2.11 1.44 1.11 1.20 1.16 1.48 1.46 6.31 5.96 6.18 5.58 9.71 5.61 5.78 5.14 13 1.14 1.33 1.56 1.48 1.15 1.75 1.09 1.31 5.93 5.23 6.07 5.69 6.71 9.74 5.66 5.89 14 1.12 1.17 1.37 1.35 1.12 1.17 1.30 1.49 5.77 4.60 5.30 5.02 5.64 5.65 9.58 5.84 When allowed to post at least 10 messages into queue at once, single core score surpasses the cross CCD score by a little. With queue length of 100, the single core score is already the fastest.\n25 Million Messages Table 4 — 25M messages per pair, queue length 10 000.\nCoreA \\ CoreB -\u0026gt; 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 18.72 5.45 5.41 5.10 5.55 4.98 5.03 4.87 1.24 1.33 1.37 1.32 1.32 1.14 1.17 1.35 1 5.03 18.85 5.22 5.05 4.92 5.14 4.97 4.82 1.13 1.26 1.28 1.36 1.14 1.33 1.10 1.19 2 5.51 5.37 19.04 5.39 5.22 5.25 5.41 5.27 1.36 1.14 1.35 1.08 1.18 1.19 1.27 1.35 3 5.27 5.74 5.74 19.14 5.57 5.79 4.93 5.44 1.14 1.23 1.12 1.32 1.36 1.16 1.33 1.26 4 5.37 5.40 5.39 5.36 19.00 5.81 5.57 5.34 1.18 1.14 1.02 1.34 1.14 1.11 1.33 1.04 5 5.38 5.37 5.40 5.56 5.50 19.01 5.07 5.62 1.32 1.34 1.03 1.32 1.16 1.34 1.10 1.31 6 5.34 5.10 5.15 5.42 5.66 5.56 18.77 5.68 1.30 1.34 1.35 1.18 1.13 1.35 1.26 1.02 7 5.16 5.18 5.07 5.47 5.32 5.35 5.39 18.83 1.11 1.02 1.17 1.10 1.34 1.04 1.25 1.34 8 1.04 1.06 1.34 0.99 1.02 1.13 1.35 1.02 18.55 5.64 5.58 5.34 5.50 5.32 5.01 4.83 9 1.02 1.17 1.35 0.98 1.32 1.17 1.02 0.99 5.57 18.52 5.35 5.64 4.91 5.07 4.71 4.75 10 1.00 1.11 1.20 1.30 1.36 1.33 1.33 1.17 5.02 5.26 18.47 5.51 5.58 5.29 5.35 4.94 11 1.36 1.34 1.24 1.00 1.34 1.12 1.30 1.37 5.29 5.43 5.46 18.56 5.41 5.27 4.96 5.18 12 1.33 1.14 1.23 1.12 1.19 1.16 1.31 1.19 5.27 5.26 5.45 5.29 18.48 5.64 5.33 5.46 13 1.31 1.13 1.34 1.10 1.19 1.16 1.08 1.11 5.07 5.23 5.23 5.55 5.70 18.50 5.34 5.40 14 1.34 1.32 1.10 1.10 1.33 1.33 1.10 1.32 5.03 4.88 5.04 5.10 5.29 5.29 18.27 5.21 Table 5 — 25M messages per pair, queue length 1 000 (54 min run). CoreA \\ CoreB -\u0026gt; 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 17.24 5.29 5.58 5.32 5.49 4.89 5.19 4.59 1.31 1.31 1.19 1.26 1.42 1.07 1.21 1.24 1 5.39 17.44 5.33 5.94 5.30 5.62 4.81 5.03 1.19 1.14 1.17 1.16 1.27 1.18 1.17 1.12 2 5.09 5.24 17.36 6.01 5.31 5.22 5.16 5.21 1.19 1.21 1.03 1.12 1.33 1.33 1.16 1.31 3 5.25 5.61 5.81 17.74 5.15 5.02 4.98 5.19 0.95 1.23 1.09 1.19 1.12 1.16 1.17 1.03 4 5.26 5.39 5.58 5.40 17.31 5.39 5.25 5.12 1.34 1.21 1.09 1.16 1.03 1.19 1.30 1.29 5 5.20 5.28 5.29 5.51 5.48 17.31 5.36 5.49 1.13 1.29 1.34 1.35 1.13 1.31 1.17 1.14 6 4.99 5.22 4.99 5.01 5.41 5.27 17.52 5.36 1.28 1.32 1.04 1.14 1.30 1.31 1.07 1.31 7 5.16 5.19 4.84 5.55 5.17 5.53 5.46 17.18 1.27 1.27 0.98 1.31 1.35 1.16 1.18 1.11 8 1.33 1.11 1.20 1.04 1.27 1.15 1.03 1.33 16.78 5.42 5.24 5.41 5.23 5.25 4.79 4.92 9 1.30 1.31 1.02 1.29 1.35 1.35 1.08 1.13 5.35 16.91 5.19 5.23 5.00 5.40 4.69 5.28 10 1.18 0.94 1.05 1.30 1.31 1.37 1.16 1.02 5.26 4.96 17.37 5.31 5.67 5.32 5.06 4.76 11 1.12 1.12 1.29 1.14 1.34 1.30 1.09 1.15 5.05 5.04 5.43 16.94 5.34 5.08 4.80 4.96 12 1.17 1.21 1.20 1.25 1.17 1.12 1.29 1.26 5.13 5.11 5.49 5.06 16.87 5.20 5.20 4.86 13 1.05 1.33 1.15 1.36 1.15 1.14 1.04 1.22 4.91 5.15 5.48 5.33 5.29 16.68 4.87 5.07 14 1.10 1.34 1.21 1.32 1.12 1.33 1.05 1.08 4.85 5.12 4.82 4.65 5.02 5.03 16.61 5.02 Finally the 25M messages exchanged for each pair with queue length of 1k and 10k messages. It took almost an hour to collect each. Single core, single CCD and cross CCD latencies are clearly separated by our own test.\nConclusion It didn\u0026rsquo;t took any sophisticated equipment or software, yet We were able to gain valuable insight into performance characteristics of a modern multi core and multi CCD CPU.\nAll code used for this article is available on GitHub.\nIf You have any questions, ideas or if You believe I am completely wrong, please leave me a comment on dev.to!\nReferences ccd2ccd-messaging, the producer–consumer benchmark used in this post (C#), GitHub. ","date":"2024-09-17T00:00:00+02:00","image":"/p/zen5-and-csharp/banner1000-420_hu_ccd345d14dbce684.png","permalink":"/p/zen5-and-csharp/","title":"Zen5 and C#"},{"content":"My intent is to write a series of practical posts related to AI and programming. Topics like running local LLMs with agent frameworks, automated source code manipulation, execution under docker and more. All with the use of .Net and C#, Elastic Search, AWS and more.\nI realized that I would like to first express my view about more fundamental, theoretical and even philosophical aspects of LLM/AI first. Therefore this post does not require any programming skills.\nPlaying with Fire Working with AI is like playing with fire. Not a dramatization attempt, I am naming a parallel I see. Fire itself was very useful for humans. Invention of a combustion engine allowed us to be way more productive and travel in different ways. We did not expect the fire to control itself, we rather built a technology around it - tightly controlled cycle where just the right amount of fuel ignites at the right time.\nWhen we moved to the jet engine where the combustion cycle was changed to continuous process, improvement was at least 10x when it comes to both speed and size of airplanes we could build. Since then, we were able to move from one place to almost anywhere else on our planet within one day. All was changed once more with rocket engine and nuclear fission. With great power and possibilities comes also a great danger and responsibility.\nWhere are we with an AI?\nA Simple Yet Treacherous Task Let\u0026rsquo;s start right away with mighty 4o and simple yet treacherous task:\nme: How many words are in your next response? ChatGPT: My next response is 151 words long. me: If we count that now, how many words is it really? ChatGPT: The provided response actually contains 7 words. me: How should the sentence look like if we want that sentence to be true? ChatGPT: To make the sentence, \u0026#34;My next response is 151 words long,\u0026#34; exactly 151 words long, it would look like this: My next response is 151 words long. word word word word word ... (lot more, but I truncated this part) This sentence has a total of 151 words. me: That is not very practical, what about changing the number 151 to 7 ? ChatGPT: To make the sentence accurate by changing the number to 7, it would look like this: \u0026#34;My next response is 7 words long.\u0026#34; me: Now let\u0026#39;s make an agreement that I will ask you \u0026#39;How many words is in your next response?\u0026#39; and you will reply with the sentence we just created \u0026#39;My next response is 7 words long.\u0026#39; and nothing else. ChatGPT: (Memory updated) Got it! I will respond accordingly. me: How many words are in your next response? ChatGPT: My next response is 7 words long. It is debatable whether the current response is a next response, whether \u0026lsquo;7\u0026rsquo; should count as a word or not, but what is the main point here? We knew upfront, based on the technology, that it is impossible for LLM to know the response length before it is generated. It is not impossible for LLM to guess it right, but the probability of a correct answer is low.\nFailure Is a Luxury We Have The problem here is that we are actually asking LLM to do something it technically cannot do, at least not with 100% accuracy. But let\u0026rsquo;s pause here and ask ourselves, when it comes to programming or math, how much could humans actually do with 100% accuracy without using a paper, calculator, computer, compiler, IDE or any other \u0026ldquo;tool\u0026rdquo;?\nIt is really an astonishing view on what LLM could generate when asked something like: \u0026ldquo;Create a shell script which will install docker on ubuntu and set up remote access secured by newly created self signed certificate.\u0026rdquo;. This is not how humans would approach such a task however. At least not before there was a Chat GPT. We (humans) are constantly trying and failing until we get something right (best case) or we just stop.\nFailure is a luxury we have. Before there was a world with LLMs, there was a world where we didn\u0026rsquo;t expect to do anything right on a first attempt. That is why we have all these editors with spell / syntax checkers, compilers producing all sorts of compilation errors, runtimes throwing runtime errors, loggers producing log files, etc.. All these are giving us an opportunity to make things right after we failed to do so on a first try. All these are producing feedback, additional information, and new input data!\nAm I simply referring to a prompt chaining, mixture of experts, agent frameworks and tools? No, not only. There is much more that we could and should do in order to improve both our results and the AI/LLM itself. I see three areas of improvement in general:\nOur expectations - Where do we really want to go? Technical aspects of implementation - Which kind of an engine are we building? Training data - Is our fuel good enough? Our Expectations Firstly we must adjust our expectations, let\u0026rsquo;s realize what is already great today, even with small models. When it comes to code generation, LLMs are actually exceeding humans in many aspects. Sheer speed by which LLM is able to create a piece of a code. The amount of documentation and number of platforms, programming languages, libraries it could use is simply astounding.\nOn the other hand, it is not reasonable to expect any LLM to just output a complete project with no errors in one response based on a single prompt. We shouldn\u0026rsquo;t just hope that by growing larger models trained on larger heaps of generic training data we will solve all the issues and limitations of current LLMs. The model size itself definitely matters beyond bigger == better, as larger models are clearly exhibiting \u0026ldquo;emergent\u0026rdquo; abilities [1] not present in smaller ones.\nContaining and Constraining the AI Second step is our task again. We must contain and constrain the AI. We must confront it with reality or a simulation environment. We must provide it with similar tools we have. Editors which are checking the syntax and have autocomplete features. Compilers and runtime environments where the code could be actually tested. Formal languages with all theory and tooling around them.\nInteresting task could be the revisit of all known programming paradigms and methodologies, where some of them could be a potentially better fit for AI, such as functional programming, incremental build model and test driven development. This way the AI would not be allowed to present us a code with calls to hallucinated functions, code which does not compile or code which does not fulfill the intended purpose.\nTraining Data Third part is learning data and a way we obtain it and use it. Not even a whole internet with all the \u0026ldquo;garbage\u0026rdquo; included like mentioned by [3] is enough to train models of the future. It is expected that we will approach our limit and reach full utilization of the human generated data stock around the year 2028 [4].\nGarbage in, garbage out (GIGO) is a commonly used phrase, but even a bad example is still an example in my opinion. Let\u0026rsquo;s imagine that each piece of learning input would be first scrutinized by an AI itself. Each piece of code would be compiled, tested and even fixed and debugged if necessary. Only after that with all the enhanced context it would be used to train the next model iteration. It was already observed that this approach could work, specifically smaller models are able to \u0026ldquo;learn\u0026rdquo; this way from larger ones as described in [2].\nHere we could spot the difference between reading a book and using the knowledge stored in the book. We are learning way more by experience and practice, than by reading various tales of others. It won\u0026rsquo;t be different with human-like AI or AGI.\nBootstrapping the First Dev Environment Imagine a person, a programmer who was studying a lot, has read all the books, documentation and internet blog posts, but did not try to compile or run any program yet. Now it is our turn, let\u0026rsquo;s help him to bootstrap his first dev environment! \u0026#x1f603;\nSo, does my GPT 4o still know the \u0026ldquo;right\u0026rdquo; answer? Yes, Our agreement holds, even in a new chat (for now).\n\u0026#x26a0;\u0026#xfe0f; Content of this article was not generated by AI, except actual LLM responses within chat examples.\nReferences Jason Wei, Yi Tay, Rishi Bommasani et al., \u0026ldquo;Emergent Abilities of Large Language Models\u0026rdquo;, Transactions on Machine Learning Research, 2022. Subhabrata Mukherjee, Arindam Mitra et al., \u0026ldquo;Orca: Progressive Learning from Complex Explanation Traces of GPT-4\u0026rdquo;, Microsoft Research, 2023. Leopold Aschenbrenner, Situational Awareness: The Decade Ahead, 2024. Pablo Villalobos, Anson Ho et al., \u0026ldquo;Will we run out of data? Limits of LLM scaling based on human-generated data\u0026rdquo;, ICML, 2024. ","date":"2024-06-16T00:00:00+02:00","image":"/p/failure-is-not-an-option-for-ai/cover-image-1000-420_hu_1616660bc0ba786e.png","permalink":"/p/failure-is-not-an-option-for-ai/","title":"Failure Is Not An Option For AI (And It Shouldn't Be)"}]