<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Peyton Walters</title><link>https://pawa.lt/</link><description>Peyton Walters</description><generator>Hugo -- gohugo.io</generator><language>en-us</language><copyright>Peyton Walters</copyright><lastBuildDate>Sat, 29 Dec 2018 00:00:00 +0000</lastBuildDate><atom:link href="https://pawa.lt/index.xml" rel="self" type="application/rss+xml"/><item><title>NYC backpacking for noobs</title><link>https://pawa.lt/posts/2025/12/nyc-backpacking-for-noobs/</link><pubDate>Wed, 24 Dec 2025 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2025/12/nyc-backpacking-for-noobs/</guid><description>I love living in the city, but my heart yearns for nature. When I first moved to the city, I got my fix by making it out to Beacon or Cold Spring, but I got tired of doing the same manicured hike for the fifth time. Backpacking has brought back my love of being in unfiltered nature - unfortunately, guides on how to do it are few and far between!</description><content type="html"><![CDATA[<p>I love living in the city, but my heart yearns for nature. When I first moved to the city, I got my fix by making it out to <a href="https://www.scenichudson.org/explore-the-valley/scenic-hudson-parks/mount-beacon-park/">Beacon</a> or <a href="https://lauraperuchi.nyc/day-trip-ideas-from-nyc-a-hiking-trail-in-cold-spring-accessible-by-train/">Cold Spring</a>, but I got tired of doing the same manicured hike for the fifth time. Backpacking has brought back my love of being in unfiltered nature - unfortunately, guides on how to do it are few and far between!</p>
<p>I&rsquo;m writing this guide in hopes to demystify backpacking out of NYC for car-less beginner backpackers like me.</p>
<h2 id="where-to-go">Where to go</h2>
<p>Without a car, you&rsquo;re bound by the limits of public transportation. Luckily, the subway and adjacent commuter rail have great access to nature. The spots mentioned here have bespoke shelters as well - in both parks, it is illegal to camp outside designated shelters or campsites.</p>
<h3 id="appalachian-trail-via-metro-north">Appalachian Trail via Metro-North</h3>
<p>From Grand Central station, you can take the Metro-North up to the <a href="https://en.wikipedia.org/wiki/Appalachian_Trail_station">Appalachian Trail</a> Metro-North station. This stop is only serviced on the weekends and holidays - on weekdays, go to Pawling and walk or Uber up to the trail.</p>
<p>As the name suggests, this drops you directly onto the Appalachian Trail. From here, the closest shelter is <a href="http://www.wikitrail.org/features/view/at/25785/telephone-pioneers-shelter">Telephone Pioneers Shelter</a>. This shelter has both a lean-to and some space to camp. Go a few miles past the shelter and you&rsquo;ll find <a href="https://hikethehudsonvalley.com/hikes/nuclear-lake/">Nuclear Lake</a>, a beautiful spot to stop and pull some water using your water filter. This is especially useful since the creek near Telephone Pioneers is frequently dry.</p>
<p>On your way back, I&rsquo;d highly recommend walking down to Pawling! It&rsquo;s a cute town in which to grab coffee or food after a long weekend.</p>
<h3 id="harriman-state-park-via-nj-transit">Harriman State Park via NJ Transit</h3>
<p>Harriman is my personal favorite due to its wide range of trails and shelters. From the train-accessible entrance, there are many forks to take. You can have a new experience every time!</p>
<p>To get to Harriman, take a train from Penn Station to Secaucus Junction and another to the Tuxedo stop. The Tuxedo line runs every day but at irregular times - make sure to plan ahead on the NJ Transit app so you know when to head out and back.</p>
<p>Once in the park, start your trek to one of the many shelters. Download the <a href="https://parks.ny.gov/sites/default/files/HarrimanTrailMap.pdf">Harriman Trail Map PDF</a>  on your phone or use AllTrails to find your way around. Note that on holidays, the main Harriman shelter areas can get quite crowded. If you have a car, this is a good opportunity to start your hike at places less accessible by train.</p>
<h2 id="gear">Gear</h2>
<p>Backpacking gear chat is centered around one thing: weight. You&rsquo;ll be carrying your pack the whole way, so even slight differences in weight matter. This is what the <a href="https://www.reddit.com/r/Ultralight/">Ultralight</a> movement aims to optimize for - we even have our own <a href="https://www.reddit.com/r/nycultralight/">r/nycultralight</a>!</p>
<p>Backpacking gear can get expensive, so it makes sense to simultaneously optimize for cost. The upshot is that the gear is quite durable and will last you for years, but the downside is the initial cost is quite high. Once you have a good setup though, you can take trip after trip for just the cost of food and transportation!</p>
<h3 id="backpack">Backpack</h3>
<p>Most backpacking packs sit in the 55-65 liter range - this balances holding enough stuff while staying light. When getting a pack, the most important thing to check is the fit. A good backpack will have waist and chest straps with a hard backplate - this allows it to sit on your back, evenly distributing weight along your back. A good fit is extremely important, so unless you&rsquo;re experienced, start at someplace like REI where you can feel out different bags.</p>
<h3 id="tent">Tent</h3>
<p>Your tent is frequently the largest item in your backpack, so it&rsquo;s important to get one that&rsquo;s right for you! There are many cheap ($50) tents on Amazon, but buy at your own risk - standard camping tents are bulkier and heavier than backpacking tents. They can get you started but will slow you down on longer hauls. I&rsquo;m currently using <a href="https://www.amazon.com/dp/B014LSDUA8">this one</a>, and it&rsquo;s absolutely painful to backpack with; it takes up more than half of my backpack alone. Upgrade coming soon.</p>
<p>There are <a href="https://www.outdoorgearlab.com/topics/camping-and-hiking/best-backpacking-tent">websites</a> that can give better tent reviews than me, but you&rsquo;re mostly optimizing for sleeping area size vs. packed size/weight. I like a two-person tent even for myself so I can lounge out. If you&rsquo;re looking to go as light as possible, some popular tents forego any inbuilt supports and use your trekking poles instead.</p>
<h3 id="sleeping">Sleeping</h3>
<p>You can use a standard sleeping bag for backpacking. I have the <a href="https://www.outdoorgearlab.com/reviews/camping-and-hiking/backpacking-sleeping-bag/rei-co-op-trailbreak-30">REI TrailBreak 30</a>, and I&rsquo;m super happy with it! Advanced users can venture into the world of <a href="https://www.cleverhiker.com/backpacking/best-backpacking-quilts/">quilts</a>, but that&rsquo;s above my paygrade.</p>
<p>To give yourself some cushion at night, you&rsquo;ll need some sleeping pad. The main options are:</p>
<ol>
<li>Air inflatable: These require some work to inflate and are prone to accident since popping the mattress is game over. However, they pack the smallest and are the most comfortable! These are my personal favorite - <a href="https://www.amazon.com/PowerLix-Sleeping-Pad-Orange-Black/dp/B00CBOFS8M">I use this cheap one</a>.</li>
<li>Self-inflating pads: These provide some cushion and autoinflation via foam, but they also seal in the air to provide a more comfortable surface. They&rsquo;re quite comfortable but come at a price. The nice ones are very expensive and the cheap ones lose their ability to inflate after being stored for too long.</li>
<li>Closed-cell foam: This is what a lot of legit backpackers go with. It&rsquo;s impossible to pop and hence deflate, and it can double as a seat cushion. It&rsquo;s also probably the least comfortable. Side sleepers beware!</li>
</ol>
<p>If you&rsquo;ll be camping in the cold, pay attention to the <a href="https://www.switchbacktravel.com/info/sleeping-pad-r-value">R value</a> of your sleeping pad. This describes how much insulation your pad will give you from the ground. Remember that you can always increase the R value of your setup by adding a closed-cell foam pad underneath your existing pad.</p>
<h3 id="misc">Misc</h3>
<p>Some important but smaller items:</p>
<ul>
<li>Water filter: allows you to cut down on your total weight by avoiding packing water. <a href="https://www.rei.com/product/204129/lifestraw-peak-series-water-filter-straw">The straws</a> attach to standard water bottles and are very easy to work with.</li>
<li>Toilet paper &amp; trowel: nuff said</li>
<li>Bear bags: absolutely critical for hiking in bear country (most of the American Northeast)</li>
<li>Regular plastic bags: you&rsquo;ll never regret another bag or two for trash or general organization.</li>
<li>Flashlight/battery pack: The night is scary without a flashlight, and your phone may die quickly. I&rsquo;m happy with <a href="https://www.amazon.com/Rechargeable-20000mAh-Waterproof-Emergency-Hurricane/dp/B08HMWDNVX/">this one off Amazon</a></li>
<li>First aid kit: Some band-aids or bandage kit are great in a pinch</li>
</ul>
<h2 id="eating">Eating</h2>
<p>For your first trip, it&rsquo;s easiest to pack cold food to avoid having to buy a stove. You can <a href="https://www.rei.com/learn/expert-advice/a-guide-to-cold-soaking-your-food.html">cold soak</a> or go with some trail classics like pasta salad, trail mix, tuna salad, or sandwiches. I always bring some sardines and crackers for quick and filling snacks.</p>
<p>That said, a camping stove and pot are a huge upgrade - they make the cold more bearable, and they allow you to cut down on your total weight. I have the <a href="https://www.outdoorgearlab.com/reviews/camping-and-hiking/backpacking-stove/soto-windmaster">Soto Windmaster</a> and the <a href="https://bikepacking.com/gear/msr-titan-kettle-review-2024-titan-cookware/">MSR Titan 900</a>. The windmaster and a fuel canister fit inside the Titan, so my whole stove setup only takes the space of my pot.</p>
<p>The easiest (and quite tasty) foods to cook are the <a href="https://mountainhouse.com/products/beef-stroganoff-pouch">pre-packaged camping meals</a>. These can get pricey and are high in sodium, so to mix things up, you could try:</p>
<ul>
<li>Rice and lentils with tinned fish over the stove - my personal favorite</li>
<li>Any instant microwave meal like mashed potatoes</li>
<li>Dehydrated meals from the internet. <a href="https://thruhikers.co/">Renee and Tim</a> have some great stuff.</li>
<li>Anything else you can dream up :)</li>
</ul>
<h2 id="some-final-advice">Some final advice</h2>
<p>Before closing out, in no particular order:</p>
<ul>
<li>Check weather and temperature before heading out. For me, the optimal times are fall and spring when the weather is 50-75°F.</li>
<li>Tell friends you&rsquo;re leaving and when you&rsquo;ll be back so they know to be suspicious if you don&rsquo;t make it back.</li>
<li>If your iPhone does not have satellite SOS capabilities, consider buying an SOS beacon.</li>
<li>Remember that you&rsquo;re a guest in nature. Practice <a href="https://lnt.org/why/7-principles/">Leave No Trace</a> so we can keep backpacking for years to come.</li>
</ul>
<p>Now go have some fun in nature with your friends!</p>
]]></content></item><item><title>Instant Pot Reference</title><link>https://pawa.lt/posts/2025/05/instant-pot-reference/</link><pubDate>Sun, 18 May 2025 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2025/05/instant-pot-reference/</guid><description>I normally cook by feel, but that doesn&amp;rsquo;t work with an Instant Pot. This is a working document for my own reference. All recipes are on high pressure. Check Amy and Jack for more detailed recipes.
Item Active Time Natural Release Time Chickpeas (unsoaked) 50 min 10 min Chickpeas (soaked) 5 min 10 min Black beans (unsoaked) 30 min 20 min Carnitas (2 inch cubes) 40 min 15 min Jasmine white rice 3 min 10 min Carnitas 45 min 15 min</description><content type="html"><![CDATA[<p>I normally cook by feel, but that doesn&rsquo;t work with an Instant Pot. This is a working document for my own reference. All recipes are on high pressure. Check <a href="https://www.pressurecookrecipes.com/">Amy and Jack</a> for more detailed recipes.</p>
<table>
<thead>
<tr>
<th>Item</th>
<th>Active Time</th>
<th>Natural Release Time</th>
</tr>
</thead>
<tbody>
<tr>
<td>Chickpeas (unsoaked)</td>
<td>50 min</td>
<td>10 min</td>
</tr>
<tr>
<td>Chickpeas (soaked)</td>
<td>5 min</td>
<td>10 min</td>
</tr>
<tr>
<td>Black beans (unsoaked)</td>
<td>30 min</td>
<td>20 min</td>
</tr>
<tr>
<td>Carnitas (2 inch cubes)</td>
<td>40 min</td>
<td>15 min</td>
</tr>
<tr>
<td>Jasmine white rice</td>
<td>3 min</td>
<td>10 min</td>
</tr>
<tr>
<td>Carnitas</td>
<td>45 min</td>
<td>15 min</td>
</tr>
</tbody>
</table>
]]></content></item><item><title>Congestion Pricing Tracker</title><link>https://pawa.lt/braindump/congestion/</link><pubDate>Sun, 05 Jan 2025 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/congestion/</guid><description>NYC finally implemented congestion pricing! There&amp;rsquo;s already a tracker to show how much it&amp;rsquo;s changed.
Credit to Emily Oster on Twitter.</description><content type="html"><![CDATA[<p>NYC finally implemented <a href="https://www.nytimes.com/live/2025/01/05/nyregion/congestion-pricing-nyc-new-jersey">congestion pricing</a>!
There&rsquo;s already a <a href="https://www.congestion-pricing-tracker.com/">tracker</a> to show how much it&rsquo;s changed.</p>
<p><img src="/img/congestion_pricing_tracker.png" alt="pricing tracker showing decrased holland tunnel traffic"></p>
<p>Credit to <a href="https://x.com/ProfEmilyOster/status/1875870736398925972">Emily Oster on Twitter</a>.</p>
]]></content></item><item><title>Using Claude Artifacts to analyze Spotify listening</title><link>https://pawa.lt/braindump/spotify-artifact/</link><pubDate>Tue, 24 Dec 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/spotify-artifact/</guid><description>My friend told me that you can request a download of all your listening activity from Spotify! I got Claude to one-shot creating an artifact to do this. You can view it right here! I&amp;rsquo;ve also got my own listening data you can use while you wait for your data to come back.
I juiced this base artifact up a little using Cursor chat, and now it looks much better! View it here.</description><content type="html"><![CDATA[<p>My friend told me that you can <a href="https://www.spotify.com/us/account/privacy/">request a download</a> of all your listening activity from Spotify! I got Claude to one-shot creating an artifact to do this. You can view it <a href="/artifacts/listening/base-activity.html">right here</a>! I&rsquo;ve also got my own <a href=https://pawa.lt/artifacts/listening/my-spotify-data.json download>listening data</a> you can use while you wait for your data to come back.</p>
<p>I juiced this base artifact up a little using <a href="https://www.cursor.com/">Cursor</a> chat, and now it looks much better! <strong><a href="/artifacts/listening/activity.html">View it here</a></strong>.</p>
<p><img src="/artifacts/listening/screenshot.png" alt="screenshot of visualizer"></p>
<p>Big thanks to <a href="https://simonwillison.net/">Simon Willison</a> for <a href="https://simonwillison.net/2024/Oct/21/claude-artifacts/">giving me the inspiration</a> to make this simple visualizer.</p>
]]></content></item><item><title>Teach systems, not facts</title><link>https://pawa.lt/posts/2024/12/teach-systems-not-facts/</link><pubDate>Fri, 20 Dec 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2024/12/teach-systems-not-facts/</guid><description>Most people believe our education system is broken, but in my experience, it&amp;rsquo;s not common that two people actually have the same gripe with the system. Complaints vary from teacher pay to content practicality to poor standardized tests. I don&amp;rsquo;t think any of those are the main problem (although they may be upstream of my personal complaint). I think the core issue is the model in which we teach kids to think: we teach them to remember facts, not to understand systems.</description><content type="html"><![CDATA[<p>Most people believe our education system is broken, but in my experience, it&rsquo;s not common that two people actually have the same gripe with the system. Complaints vary from teacher pay to content practicality to poor standardized tests. I don&rsquo;t think any of those are the main problem (although they may be upstream of my personal complaint). I think the core issue is the model in which we teach kids to think: we teach them to remember facts, not to understand systems.</p>
<p>My best example of this is the <a href="https://www.kumon.com/">Kumon</a> method of teaching. Kumon works by having students read some material about their next unit. After reading this material, they do worksheets in the classroom and at home. These worksheets are graded when they come in for their next session, and if they&rsquo;re doing well enough, they get the opportunity to take a test and graduate to the next level. Students will learn arithmetic at a very young age and become proficient quite quickly. This makes for cosmetically impressive results - I&rsquo;ve never seen anyone do multiplication faster than a 7-year-old who&rsquo;s been doing Kumon for years.</p>
<p>I think these results are a mirage. While students are able to progress through the units, they do it by learning tricks. They learn the <a href="https://study.com/academy/lesson/dividing-compound-fractions.html">keep-change-flip</a> method of dividing fractions, or they learn to solve quadratic equations by memorizing the <a href="https://letmegooglethat.com/?q=quadratic+formula">quadratic formula</a>. These tricks will get you the right answer, but they don&rsquo;t help you build any deeper intuition about what&rsquo;s going on. Learning the quadratic formula won&rsquo;t help you understand differentiation better, but <a href="https://www.mathsisfun.com/algebra/completing-square.html">completing the square</a> will! The keep-change-flip method won&rsquo;t help you understand negative exponents, but splitting up the fractions into their component parts will. In both cases, the &ldquo;trick&rdquo; will get you to the right answer reliably and is fast to learn. It won&rsquo;t, however, help you generalize onto future situations.</p>
<p>I&rsquo;m picking on Kumon, but I don&rsquo;t think the problem is exclusive to that system; they&rsquo;re just the worst offenders that I know of (I was a Kumon tutor in a past life). I&rsquo;ve had plenty of math teachers try to teach through these short-sighted tricks. I don&rsquo;t blame them - their explicit reward system is only tied to standardized test results and students&rsquo; grades. In this system, I would also try to shuttle along kids so I could get along with my job. It&rsquo;s only the best teachers, the ones truly in it for the love of the game, that put in the effort to teach differently. I&rsquo;ve had a <a href="https://informationtechnology.henricoschools.us/page/teachers/">few of these</a> kinds of teachers, and they changed my life.</p>
<p><strong>The way we teach kids how to learn in school has ripple effects on the rest of their lives.</strong> In fact, I think this problem only rears its head in higher-level education or in the real world. In a real job, these tricks don&rsquo;t exist in the same way: there&rsquo;s no textbook with which to learn shortcuts to the &ldquo;right answer&rdquo; if such a provably right answer even exists. Instead, you&rsquo;re forced to take in your surroundings and build a mental model for how your domain works. Building a strong understanding of your domain helps when doing routine tasks, and it&rsquo;s <em>literally required</em> when trying to build something novel.</p>
<p>I&rsquo;m particularly interested in this because so much of my &ldquo;education&rdquo; has happened outside the classroom. I&rsquo;ve learned primarily through <a href="/posts/2019/07/caplance-development-update-3/">side projects</a> and <a href="/braindump/pushing-yourself/">learning on the job</a>. In either case, I didn&rsquo;t read a book and memorize some tricks; nor did I just try to accomplish the task in the most expedient way. Instead, I&rsquo;ve tried to focus on the process of <a href="/braindump/dag-building/">building my DAG</a>. By focusing on building a wider understanding of the system, I&rsquo;ve been able to generalize to new situations and build novel solutions. For me, this deeper understanding is also more durable: when I can slot information into my DAG, it sticks with me near-permanently.</p>
<p>While it&rsquo;d be nice to neatly wrap up this post with a suggestion on how to incentivize this teaching style, I honestly don&rsquo;t know how. Pedagogy reforms like Common Core have largely <a href="https://www.brookings.edu/articles/why-common-core-failed/">been a disaster</a> - this style of instruction requires a deep understanding of the material and the teaching ability to execute on this understanding. Even if we had a fleet of teachers all of whom were capable of teaching like this, they aren&rsquo;t incentivized to. <strong>The benefits of this way of thinking play out over years</strong>, so you can&rsquo;t reward this behavior on a per-teacher basis. It would never show up on a standardized test.</p>
<p>It may be that we need new <a href="/braindump/scalable-tutors/">interactive styles of instruction</a> to unlock this, but these beliefs would need to be embedded into those systems. All I know is that when I&rsquo;m teaching, I&rsquo;ll continue to encourage my students to think about the wider picture. I hope that you will too.</p>
]]></content></item><item><title>Finding My Voice</title><link>https://pawa.lt/posts/2024/12/finding-my-voice/</link><pubDate>Sun, 15 Dec 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2024/12/finding-my-voice/</guid><description>Growing up, I thought of myself as a bad writer. That&amp;rsquo;s not to say I couldn&amp;rsquo;t write - I could always pump out a 5 paragraph essay. 1 intro, 3 body, 1 conclusion. 5 sentences each. It felt, though, like some people had something I didn&amp;rsquo;t. I could follow the formula and make a point, but I struggled to go beyond that and break the mold. I just couldn&amp;rsquo;t find my voice.</description><content type="html"><![CDATA[<p>Growing up, I thought of myself as a bad writer. That&rsquo;s not to say I couldn&rsquo;t write - I could always pump out a 5 paragraph essay. 1 intro, 3 body, 1 conclusion. 5 sentences each. It felt, though, like some people had something I didn&rsquo;t. I could follow the formula and make a point, but I struggled to go beyond that and break the mold. I just couldn&rsquo;t find my voice.</p>
<p>This was infuriating because I&rsquo;ve never had trouble finding my voice orally. My whole life I&rsquo;ve been speaking, teaching, and debating with passion. These experiences always felt different to me: when writing I felt friction just arranging my words onto the page, but when speaking, I felt a flow that allowed me to say exactly what I meant <em>exactly</em> how I wanted. At the time, I saw writing as an immutable weakness of mine. Some people are good at some things, some people are good at other things.</p>
<p>This feeling changed for me when I started to write about things I enjoyed. I have a distinct memory of writing a long essay about Edward Snowden at some summer camp - I turned it in expecting for it to be ripped apart, but I was met with praise instead. I built a part of my identity around this weakness, accepting that some people are good at some things while some are good at others. While there&rsquo;s some truth here, this way of viewing myself was artificially limiting.</p>
<p>And so I wrote about why to use <a href="https://fishshell.com/">fish</a> not <a href="https://www.gnu.org/software/bash/">bash</a>; about why I should buy a <a href="https://www.kickstarter.com/projects/getpebble/pebble-e-paper-watch-for-iphone-and-android">Pebble</a>; about <a href="/posts/2018/01/vpls-with-openbsd/">random bits of network engineering</a>. This moved me in the right direction, yet it wasn&rsquo;t enough. Even though I started to write about things that lit me up, I still had expectations I was holding onto. Whether it was my boss, teacher, or some imaginary hyper-critical netizen, I was worried that someone more skilled than me would read my work and look down on me. For this reason, I wrote almost exclusively technical posts. Tech is something I&rsquo;ve always been confident in; it allowed me to write without fear of judgement.</p>
<p>I have so much more to say than just explainers about tech! It&rsquo;s only been through writing <a href="/braindump/">braindumps</a> that I&rsquo;ve felt comfortable getting those words out, though. By giving myself a space to write with no expectations, I&rsquo;ve written some of my most proud words. I&rsquo;ve had the chance to profess <a href="/braindump/dag-building/">my model of teaching</a>, thank the <a href="/braindump/pushing-yourself/">best mentor I&rsquo;ve ever had</a>, reflect on <a href="/braindump/friends/">friendship</a>, and much more. In letting go of my fear of being perceived, I&rsquo;ve finally found my voice.</p>
<p>But now that I&rsquo;ve found it, what will I say?</p>
]]></content></item><item><title>Prototyping</title><link>https://pawa.lt/braindump/prototyping/</link><pubDate>Sun, 13 Oct 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/prototyping/</guid><description>I do basically all of my interesting work through prototypes. In particular, when trying to make a complex change or build out a new system, I start by building out a shit-quality MVP that simply proves the thing I want to do is possible. The code that I write in this process will never see the light of day, but it&amp;rsquo;s the foundation for the rest of my work.
This style of work is foreign in a lot of the industry where design docs are the primary planning vehicle.</description><content type="html"><![CDATA[<p>I do basically all of my interesting work through prototypes. In particular, when trying to make a complex change or build out a new system, I start by building out a shit-quality MVP that simply proves the thing I want to do is possible. The code that I write in this process will never see the light of day, but it&rsquo;s the foundation for the rest of my work.</p>
<p>This style of work is foreign in a lot of the industry where design docs are the primary planning vehicle. Ideally, design docs let you gather feedback and de-risk the execution of an idea before you start writing code. This is not my experience. In particular, when executing refactors, the devil is in the details - small edges in how an API is used can change large parts of its implementation. With design docs, it&rsquo;s very difficult to foresee these edges; a working prototype can teach you much more.</p>
<p>This isn&rsquo;t to say that design docs aren&rsquo;t valuable! After building my prototype, I&rsquo;ll reflect on it and write a design doc so that others are aware of my plans and can give any guiding feedback. I do find, though, that having a working prototype tends to end a lot of debate. We can get lost in the &ldquo;is this possible&rdquo; weeds; the point of a prototype is to remove this ambiguity.</p>
<p>After making my low-quality prototype and getting light consensus, I PR in my prototype bit-by-bit. In this pass, I&rsquo;ll write nice docs, write tests, and smooth out the hairy edges of my prototype. At this point, I&rsquo;m effectively doing a rewrite of my previous implementation, so I get to make all the decisions I wish I had made in the first place! I think this results in higher-quality code than writing it cold off a design doc.</p>
<p>This approach is not reasonable at all scales. At some point, it&rsquo;s infeasible to write a prototype, and you&rsquo;re better off designing with trusted co-authors. This is especially true in larger orgs with cross-functional dependencies. That said, I think you can push prototypes far further than people expect with mocking and getting your hands dirty in others&rsquo; codebases.</p>
<p>This also relies on a particular skill: being able to write bad code very quickly. The core element of execution is the ability to write a prototype quickly enough where you&rsquo;re not slowing down your total velocity. Luckily, I&rsquo;m great at writing bad code.</p>
]]></content></item><item><title>You can do so much</title><link>https://pawa.lt/braindump/pushing-yourself/</link><pubDate>Sun, 13 Oct 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/pushing-yourself/</guid><description>The greatest lesson I ever learned was that I could do so much more than I thought. I was taught this by Jon, my boss at a small datacenter I worked at in high school. Jon constantly pushed me out of my comfort zone, and I&amp;rsquo;ll forever be grateful for the lessons he taught me.
As a 16-year-old being given the opportunity to do network engineering part-time, I was just grateful to have a job that I enjoyed.</description><content type="html"><![CDATA[<p>The greatest lesson I ever learned was that I could do so much more than I thought. I was taught this by Jon, my boss at a <a href="https://richweb.com/">small datacenter</a> I worked at in high school. Jon constantly pushed me out of my comfort zone, and I&rsquo;ll forever be grateful for the lessons he taught me.</p>
<p>As a 16-year-old being given the opportunity to do network engineering part-time, I was just grateful to have a job that I enjoyed. I quickly picked up Linux and networking basics, but it was soon on to the next thing. Once I picked up one skill, I was given a new assignment with higher complexity. This process of laddering never ceased.</p>
<p>Most of the time, I didn&rsquo;t believe that I could do the next task, but Jon had an almost maniacal belief in me. This belief gave me no choice to but to push myself into learning and getting the job done. I&rsquo;d be at client sites with broken networks and no choice other than to fix the issue. I&rsquo;d be staring down borked libvirt configurations with no choice but to figure out the right hardware options. Over time, I grew to love this feeling - the feeling of never knowing the answer, always having to work to figure it out.</p>
<p>I now crave this experience. I want to constantly push myself into more difficult work, and I want to excel in that work. In part, this is out of a love for the process, but it&rsquo;s also out of a feeling of obligation. I feel that I&rsquo;ve seen through a one-way door; I can&rsquo;t go back to not pushing myself when I know how much more I can accomplish.</p>
<p>The challenge is that nobody will ever push me as hard as Jon did. I&rsquo;ll never have someone with that maniacal belief in my abilities, so in his place, I have to push myself. I hope that I can one day pass on the favor to someone else.</p>
]]></content></item><item><title>Teaching is not broadcasting</title><link>https://pawa.lt/braindump/scalable-tutors/</link><pubDate>Sun, 19 May 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/scalable-tutors/</guid><description>I&amp;rsquo;ve spent a lot of my life getting good at broadcast-style instruction, specifically lecturing and technical writing. These are both difficult skills because you have to learn to build DAGs. When teaching, you have to understand the delta between your understanding and your student&amp;rsquo;s understanding. When doing broadcast-style instruction, you do this for many students concurrently.
It&amp;rsquo;s not only that this is hard but also that it forces you to teach suboptimally to each student.</description><content type="html"><![CDATA[<p>I&rsquo;ve spent a lot of my life getting good at broadcast-style instruction, specifically lecturing and technical writing. These are both difficult skills because you have to learn to <a href="/braindump/dag-building/">build DAGs</a>. When teaching, you have to understand the delta between your understanding and your student&rsquo;s understanding. When doing broadcast-style instruction, you do this for many students concurrently.</p>
<p>It&rsquo;s not only that this is hard but also that it forces you to teach suboptimally to each student. When building a single piece of curriculum to teach everyone, you have to build up from your students&rsquo; lowest point of shared knowledge. This means you&rsquo;ll go too slowly for the high-performers, and you&rsquo;ll inevitably leave some students in the dust. In school we try to control for this with prerequisites and waitlist screening forms, but these tools are far too coarse-grained to be useful.</p>
<p>Written materials like textbooks solve some of this since students can move at their own pace, but they don&rsquo;t solve the interactivity problem. The only way to learn at max difficulty &amp; speed is to do so interactively. When you learn something, you must think hard about it and test your understanding by asking probing questions. Textbooks can teach you the information, but they can&rsquo;t help you test the boundaries of your understanding.</p>
<p>We use broadcasting because it&rsquo;s economically scalable. One teacher can broadcast to 30 students. One textbook can reach millions. In a perfect world, we&rsquo;d be able to give each student their own PhD-level tutor who can tailor instruction to exactly their knowledge and skill level. Sadly, this is cost-prohibitive.</p>
<p>I have high hopes for AI tutors as a way to solve this problem. Seeding an LLM with high quality data will allow users to interactively learn about the topic of their choice in a high-bandwidth, personalized, and social way. Hopefully, my kids will be much smarter than I am.</p>
]]></content></item><item><title>Half My Life</title><link>https://pawa.lt/posts/2024/02/half-my-life/</link><pubDate>Fri, 16 Feb 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2024/02/half-my-life/</guid><description>I turn 24 today which means I&amp;rsquo;ve spent half my life writing code. When I started, I didn&amp;rsquo;t think this was the yardstick I&amp;rsquo;d be measuring my life against, but sometimes the cookie crumbles in ways you don&amp;rsquo;t foresee.
When I think about that 7th grader booting up the Codecademy Javascript course, all I can do is remember the joy of those early days. I remember building my choose your own adventure batman game.</description><content type="html"><![CDATA[<p>I turn 24 today which means I&rsquo;ve spent half my life writing code. When I started, I didn&rsquo;t think this was the yardstick I&rsquo;d be measuring my life against, but sometimes the cookie crumbles in ways you don&rsquo;t foresee.</p>
<p>When I think about that 7th grader booting up the Codecademy Javascript course, all I can do is remember the joy of those early days. I remember building my choose your own adventure batman game. I remember making a silly little paint application. I remember building Tennis Scorekeeper+ (the plus was so people would think it was the pro version). Most of all, I remember the complete joy of writing code with no strings attached. Just in it for the love of the game.</p>
<p>Being in it &ldquo;for the love of the game&rdquo; is still the core of my experience. There&rsquo;s a lot of noise out there about the next wave to jump on. There&rsquo;s a lot of ladder climbing that will undeniably increase your TC or land you a job at <a href="https://fortelabs.com/blog/theory-of-constraints-102-local-optima/"><redacted></a>. While I don&rsquo;t have a moral issue with any of that (get that bag), it&rsquo;s just never been motivating for me. I&rsquo;ve always just wanted to work on problems I find interesting with people I like. And maybe do some good for the world.</p>
<p>I want to hold onto this for as long as I possibly can. I see far too many people jaded about their jobs, playing the corporate game so they can get that new apartment. While playing the ladder will lead to short-term benefits, I genuinely believe that the long arc of time favors those who are just trying to do their best for the sake of doing their best.</p>
<p>The last 12 years have been filled with surprise after surprise. Software has brought me success that I could have never imagined. It&rsquo;s taken me far from home to a place I thought I would despise but instead love. It&rsquo;s brought me my closest friends in life.</p>
<p>The road here has been long, and it&rsquo;s been paved with the goodwill of so many around me. The support of my family; the camaraderie of my friends; the wisdom of my mentors. I&rsquo;m grateful for all of it, and I look forward to continuing to pave it for many years.</p>
<p>I just wish my dad was here to see it.</p>
]]></content></item><item><title>Reusing Nix config across modules</title><link>https://pawa.lt/braindump/dry-nix/</link><pubDate>Sat, 20 Jan 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/dry-nix/</guid><description>Recently, I wanted to share my Syncthing config across multiple NixOS hosts, and I couldn&amp;rsquo;t find any good information on it online. I hope this helps someone. I have the following config that&amp;rsquo;s parameterized only by the username of the syncthing user - everything else stays the same:
services.syncthing = { enable = true; user = user; configDir = &amp;#34;/home/${user}/.config/syncthing&amp;#34;; dataDir = &amp;#34;/home/${user}/.config/syncthing/db&amp;#34;; overrideDevices = true; overrideFolders = true; settings = { devices = { # map from device name to ID }; folders = { # generic sync folder &amp;#34;cccjw-5fcyz&amp;#34; = { path = &amp;#34;/home/${user}/sync&amp;#34;; devices = [ # list of device names to share with ]; }; }; }; }; To modularize this, I created a Nix function with the user parameter:</description><content type="html"><![CDATA[<p>Recently, I wanted to share my <a href="https://syncthing.net/">Syncthing</a> config across multiple NixOS hosts, and I couldn&rsquo;t find any good information on it online. I hope this helps someone. I have the following config that&rsquo;s parameterized only by the username of the syncthing user - everything else stays the same:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-nix" data-lang="nix">services<span style="color:#f92672">.</span>syncthing <span style="color:#960050;background-color:#1e0010">=</span> {
  enable <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span>;
  user <span style="color:#f92672">=</span> user;
  configDir <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;/home/</span><span style="color:#e6db74">${</span>user<span style="color:#e6db74">}</span><span style="color:#e6db74">/.config/syncthing&#34;</span>;
  dataDir <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;/home/</span><span style="color:#e6db74">${</span>user<span style="color:#e6db74">}</span><span style="color:#e6db74">/.config/syncthing/db&#34;</span>;

  overrideDevices <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span>;
  overrideFolders <span style="color:#f92672">=</span> <span style="color:#66d9ef">true</span>;

  settings <span style="color:#f92672">=</span> {
    devices <span style="color:#f92672">=</span> {
      <span style="color:#75715e"># map from device name to ID</span>
    };

    folders <span style="color:#f92672">=</span> {
      <span style="color:#75715e"># generic sync folder</span>
      <span style="color:#e6db74">&#34;cccjw-5fcyz&#34;</span> <span style="color:#f92672">=</span> {
        path <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;/home/</span><span style="color:#e6db74">${</span>user<span style="color:#e6db74">}</span><span style="color:#e6db74">/sync&#34;</span>;
        devices <span style="color:#f92672">=</span> [
          <span style="color:#75715e"># list of device names to share with</span>
        ];
      };
    };
  };
};
</code></pre></div><p>To modularize this, I created a <a href="https://nixos.org/guides/nix-pills/functions-and-imports">Nix function</a> with the <code>user</code> parameter:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-nix" data-lang="nix"><span style="color:#75715e"># ./custom/syncthing.nix</span>
{ user }: {
  <span style="color:#75715e"># same config from above</span>
  <span style="color:#f92672">...</span>
}
</code></pre></div><p>Then, in my individual NixOS configuration files, I can just import it in my imports line and pass in the right param:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-nix" data-lang="nix"><span style="color:#75715e"># ./systems/asahi/default.nix</span>
{ config<span style="color:#f92672">,</span> pkgs<span style="color:#f92672">,</span> <span style="color:#f92672">...</span> }:

{
  imports <span style="color:#f92672">=</span> [
    <span style="color:#e6db74">./asahi-hardwarecfg.nix</span>
    ( <span style="color:#f92672">import</span> <span style="color:#e6db74">../../custom/syncthing.nix</span> { user <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;peyton&#34;</span>; } )
  ];
<span style="color:#f92672">...</span>
</code></pre></div><p>And voilà! I&rsquo;ve now DRYed up my configuration files. Here&rsquo;s the <a href="https://github.com/pawalt/setup/blob/334fa8d20d15716bc72fc8e66bd37049b217e0cf/custom/syncthing.nix#L1">config for the curious</a>.</p>
]]></content></item><item><title>Teaching is DAG-building</title><link>https://pawa.lt/braindump/dag-building/</link><pubDate>Sat, 20 Jan 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/dag-building/</guid><description>In order to understand a given concept, there are a set of &amp;ldquo;dependent concepts&amp;rdquo; that are required. Learning multiplication requires an understanding of addition. Learning how to cook a steak requires an understanding of temperature control. Learning Kubernetes requires an understanding of containers.
In my head, I model these relationships as a DAG. Each node is a concept, and concepts have arrows to the other concepts that depend on them. In order to fully understand one node, you must first explore all the nodes that point to it.</description><content type="html"><![CDATA[<p>In order to understand a given concept, there are a set of &ldquo;dependent concepts&rdquo; that are required. Learning multiplication requires an understanding of addition. Learning how to cook a steak requires an understanding of temperature control. Learning Kubernetes requires an understanding of containers.</p>
<p><img src="/img/kube_horizontal.svg" alt="kubernetes learning DAG"></p>
<p>In my head, I model these relationships as a <a href="https://en.wikipedia.org/wiki/Directed_acyclic_graph">DAG</a>. Each node is a concept, and concepts have arrows to the other concepts that depend on them. In order to fully understand one node, you must first explore all the nodes that point to it. You have to bottom out the recursion at some point by taking some facts as axiomatic, but IMO, the further you can go down the better.</p>
<p>As a teacher, your job is to make this DAG exist in the heads of your students. Doing this requires an understanding of two things: how you&rsquo;ve built your DAG and the current state of your students&rsquo; DAGs. These are both deceptively difficult things to understand.</p>
<p>The obvious one comes first: it&rsquo;s impossible to know what your students&rsquo; level of understanding is. The best you can do is try to guess based on past experience and filter admits to your class based on some criteria. But no matter what you do, different students will come in at different levels of understanding, and students will have gaps in their knowledge you never predicted.</p>
<p>The harder part for me, though, is grokking how my own understanding is built. When you&rsquo;re deep in a subject, you&rsquo;ve developed an intuitive understanding so deep that you don&rsquo;t even realize it&rsquo;s there. When teaching Kubernetes, you preach the wonders of its networking model only to realize its value is lost on kids who&rsquo;ve never dealt with HA systems. It&rsquo;s not that you expected them to understand high availability - it&rsquo;s just that you forgot it was an integral part of your thought process.</p>
<p><img src="/img/kube_hidden_horizontal.svg" alt="kubernetes learning DAG with hidden horizontal availability"></p>
<p>It&rsquo;s for all these reasons that I think <strong>most professors are very bad teachers</strong>. High-level research staff are so far away from undergrads that they can&rsquo;t begin to empathize with their mindset. And they certainly can&rsquo;t introspect enough to fully understand how their DAG is built.</p>
<p>Hopefully in the future we&rsquo;ll see more <a href="https://www.cis.upenn.edu/~cis19x/">classes taught by students</a>, <a href="https://www.seas.upenn.edu/~cis120/24sp/staff/">TAs that are undergrads</a>, and tenure structures that reward the <a href="https://www.cis.upenn.edu/~swapneel/">truly exceptional instructors</a>.</p>
]]></content></item><item><title>The unreasonable effectiveness of having friends</title><link>https://pawa.lt/braindump/friends/</link><pubDate>Tue, 02 Jan 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/friends/</guid><description>The world of programming content is a shitshow. From courses promising to get you jobs to influencers preaching the glory of FAANG, it&amp;rsquo;s so hard to find anything worth paying attention to.
The hard part here is that I think it&amp;rsquo;s super important to stay up to date on things! It&amp;rsquo;s probably a good idea to keep track of the LLM space, but it&amp;rsquo;s so hard to avoid the flood of grifters.</description><content type="html"><![CDATA[<p>The world of programming content is a shitshow. From <a href="https://www.algoexpert.io/product">courses</a> promising to get you jobs to <a href="https://www.youtube.com/watch?v=jmONbYqYaRk">influencers</a> preaching the glory of FAANG, it&rsquo;s so hard to find anything worth paying attention to.</p>
<p>The hard part here is that I think it&rsquo;s super important to stay up to date on things! It&rsquo;s probably a good idea to keep track of the LLM space, but it&rsquo;s so hard to avoid the flood of <a href="https://twitter.com/itsPaulAi">grifters</a>. Somehow I&rsquo;ve managed to avoid getting sucked into these pits, but I&rsquo;ve never been sure why.</p>
<p>It took way too long to come to this realization, but for me, it&rsquo;s all about the people I&rsquo;m around. My friends are the most talent-dense group of people I&rsquo;ve met in my life, and I learn so much just by being around them. Whatever online gurus can offer, you can get 1000x that just by being around people who love the game just as much as you do.</p>
<p>For me, this usually just means seeing what lands in the group chat every day. Whether it&rsquo;s the <a href="https://www.theverge.com/2023/11/17/23965982/openai-ceo-sam-altman-fired">corporate drama of the day</a> or someone going <a href="https://github.com/davish/website-astro/blob/a3874dd2770507f4256630a2dff798b88f9208af/flake.nix">way too deep</a> on their website&rsquo;s build system, I wake up every day looking forward to learning something new.</p>
<p>And all they&rsquo;ve taught me doesn&rsquo;t hold a candle to how they&rsquo;ve been there for me ❤️</p>
]]></content></item><item><title>You gotta love this thing</title><link>https://pawa.lt/braindump/love-this-thing/</link><pubDate>Tue, 02 Jan 2024 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/love-this-thing/</guid><description>I came across one of my favorite quotes ever sampled in a song, and I had to share it here.
If you&amp;rsquo;re starting to trying to be a musician or artist—something like that—because you wanna make money, because you wanna do a job, that&amp;rsquo;s- that&amp;rsquo;s the wrong way. You have to do this because you love it. And it doesn&amp;rsquo;t matter if you broke, you still gon&amp;rsquo; do it. I mean, I go out to jam sessions, and I play regardless of whether I&amp;rsquo;m getting a check or not.</description><content type="html"><![CDATA[<p>I came across one of my favorite quotes ever sampled in a song, and I had to share it here.</p>
<blockquote>
<p>If you&rsquo;re starting to trying to be a musician or artist—something like that—because you wanna make money, because you wanna do a job, that&rsquo;s- that&rsquo;s the wrong way. You have to do this because you love it. And it doesn&rsquo;t matter if you broke, you still gon&rsquo; do it. I mean, I go out to jam sessions, and I play regardless of whether I&rsquo;m getting a check or not. It&rsquo;s-, it&rsquo;s about whether I, uh—you have to love this thing, man! You have to love it and breathe it and—It&rsquo;s your morning coffee. It&rsquo;s your food. That&rsquo;s why you become an artist.</p>
</blockquote>
<p>Interview with Roy Hargrove: <a href="https://www.youtube.com/watch?v=mIaONoYeDh4">https://www.youtube.com/watch?v=mIaONoYeDh4</a></p>
<p>The greatest intro to ever hit an album: <a href="https://open.spotify.com/track/1ZwejHvd2KmKCWHn9HpAEw">https://open.spotify.com/track/1ZwejHvd2KmKCWHn9HpAEw</a></p>
]]></content></item><item><title>Why are your models so big?</title><link>https://pawa.lt/braindump/tiny-models/</link><pubDate>Tue, 26 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/tiny-models/</guid><description>I don&amp;rsquo;t understand why today&amp;rsquo;s LLMs are so large. Some of the smallest models getting coverage sit at 2.7B parameters, but even this seems pretty big to me.
If you need generalizability, I totally get it. Things like chat applications require a high level of semantic awareness, and the model has to respond in a manner that&amp;rsquo;s convincing enough to its users. In cases where you want the LLM to produce something human-like, it makes sense that the brains would need to be a little juiced up.</description><content type="html"><![CDATA[<p>I don&rsquo;t understand why today&rsquo;s LLMs are so large. Some of the smallest models getting coverage <a href="https://huggingface.co/microsoft/phi-2">sit at 2.7B parameters</a>, but even this seems pretty big to me.</p>
<p>If you need generalizability, I totally get it. Things like chat applications require a high level of semantic awareness, and the model has to respond in a manner that&rsquo;s convincing enough to its users. In cases where you want the LLM to produce something human-like, it makes sense that the brains would need to be a little juiced up.</p>
<p>That said, LLMs are a <a href="https://pawa.lt/posts/2023/08/chatting-with-my-recipes/">whole lot more</a> than just bots we can chat with. There are some domains that have a tightly-scoped set of inputs and require the model to always respond in a similar way. Something like SQL autocomplete is a good example - completing a single SQL query requires a very small context window, and it requires no generalized knowledge of the English language.  Structured extraction is similar: you don&rsquo;t need 2.7B parameters to go from <code>remind me at 7pm to walk the dog</code> to <code>{ &quot;time&quot;: &quot;7pm&quot;, &quot;reminder&quot;: &quot;walk the dog&quot; }</code>.</p>
<p>I say all this because <a href="https://www.forbes.com/sites/craigsmith/2023/09/08/what-large-models-cost-you--there-is-no-free-ai-lunch/">inference is expensive</a>. Not only is it expensive in terms of raw compute - maintaining the infrastructure required to run models also gets pretty complicated. You either end up shelling out money for in-house talent or paying some provider to do the inference for you. In either case, you&rsquo;re paying big money every time a user types <code>remind me to eat a sandwich</code>.</p>
<p>I think the future will be full of much smaller models trained to do specific tasks. Some tooling to build these <a href="https://github.com/karpathy/llama2.c">already exists</a>, and you can even <a href="https://github.com/xenova/transformers.js">run them in the browser</a>. This mode of deployment is inspiring to me, and I&rsquo;m optimistic about a future where <a href="https://ggerganov.com/llama2.c/">15M params is all you need</a>.</p>
]]></content></item><item><title>NixOS on the desktop is a pain</title><link>https://pawa.lt/braindump/nixos-is-a-pain/</link><pubDate>Sun, 24 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/nixos-is-a-pain/</guid><description>Let me preface by saying, NixOS running on Asahi is the best Linux experience I&amp;rsquo;ve ever had. Every time I&amp;rsquo;ve run a Linux distro, I always end up hand-jamming config everywhere and forgetting about it. Over time this config builds up, and after ~1 year, a full wipe is in order. NixOS has basically solved this problem for me - all my configuration is versioned as code, so skew is impossible.</description><content type="html"><![CDATA[<p>Let me preface by saying, <a href="https://nixos.org/">NixOS</a> running on <a href="https://asahilinux.org/">Asahi</a> is the best Linux experience I&rsquo;ve ever had. Every time I&rsquo;ve run a Linux distro, I always end up hand-jamming config everywhere and forgetting about it. Over time this config builds up, and after ~1 year, a full wipe is in order. NixOS has basically solved this problem for me - all my configuration is versioned as code, so skew is impossible.</p>
<p>But this experience is not for the faint of heart.</p>
<p>When the package (and version) you need is inside of <code>nixpkgs</code>, you&rsquo;re living the good life. You can just add the package name to your package list, hit a <code>switch</code>, and start using the package.</p>
<p>When you need an older version of a package, you end up trolling through git history and <a href="https://github.com/pawalt/personal-site/blob/ef42d120310b054d85ace54f80d07a3fcfc9226a/flake.nix#L4">locking nixpkgs</a> to the revision containing the version you want. If you want a newer version, you might have to update the derivation yourself and PR it in!</p>
<p>You can avoid some of these headaches by <a href="https://github.com/pawalt/setup/blob/c46fedbdbbd71cfcad6fac0a66661015a2916277/overlays/ollama.nix#L14">writing an overlay</a>, but depending on how cleanly the nixpkgs derivation is written, this can range from &ldquo;wow that was easy&rdquo; to &ldquo;why have I spent 10 hours on this&rdquo;.</p>
<p>And in the apocalyptic case where you can&rsquo;t get a derivation written, you might need to <a href="https://nixos.org/manual/nixpkgs/stable/#sec-fhs-environments">build an FHS environment</a>. But even if you do this, programs outside the FHS environment can&rsquo;t use the FHS tooling. This leaves you either putting everything inside this FHS environment or gitting gud at writing derivations.</p>
<p>I&rsquo;ll keep using NixOS, but boy do I miss the days of <code>curl https://cooltool.com/install.sh | sh</code>.</p>
]]></content></item><item><title>Asahi Linux</title><link>https://pawa.lt/braindump/asahi/</link><pubDate>Sat, 23 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/asahi/</guid><description>The Asahi project is a set of projects &amp;amp; people focused on building out support for ARM Mac devices in the Linux kernel.
Asahi can technically work with any distro, but the &amp;ldquo;officially supported&amp;rdquo; (insane levels of air quotes here) is the Fedora Asahi remix. I ran this for a while and had a really excellent experience. It was the smoothest out-of-the-box Linux experience I&amp;rsquo;ve ever had. Low bar there, but we&amp;rsquo;ll take what we can get.</description><content type="html"><![CDATA[<p>The <a href="https://asahilinux.org/about/">Asahi project</a>  is a set of projects &amp; people focused on building out support for ARM Mac devices in the Linux kernel.</p>
<p>Asahi can technically work with any distro, but the &ldquo;officially supported&rdquo; (insane levels of air quotes here) is the <a href="https://asahilinux.org/fedora/">Fedora Asahi remix</a>. I ran this for a while and had a really excellent experience. It was the smoothest out-of-the-box Linux experience I&rsquo;ve ever had. Low bar there, but we&rsquo;ll take what we can get.</p>
<p>Asahi can work with NixOS, and I&rsquo;ve had a pretty good experience running this setup. There are some <a href="https://github.com/tpwrules/nixos-apple-silicon">prebuilt NixOS modules</a> to make this easier.</p>
<p><a href="https://www.reddit.com/r/AsahiLinux/">The Asahi subreddit</a> is a good place to go for more info and support.</p>
]]></content></item><item><title>Getting open ports on macOS</title><link>https://pawa.lt/braindump/open-ports/</link><pubDate>Sat, 23 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/open-ports/</guid><description>I can never remember how to do this thing. I grew up on ss. So for posterity:
programs.zsh.shellAliases = { ss = &amp;#34;sudo lsof -PiTCP -sTCP:LISTEN&amp;#34;; };</description><content type="html"><![CDATA[<p>I can never remember how to do this thing. I grew up on <a href="https://man7.org/linux/man-pages/man8/ss.8.html">ss</a>. So for posterity:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-nix" data-lang="nix">programs<span style="color:#f92672">.</span>zsh<span style="color:#f92672">.</span>shellAliases <span style="color:#960050;background-color:#1e0010">=</span> {
  ss <span style="color:#f92672">=</span> <span style="color:#e6db74">&#34;sudo lsof -PiTCP -sTCP:LISTEN&#34;</span>;
};
</code></pre></div>]]></content></item><item><title>Nix</title><link>https://pawa.lt/braindump/nix/</link><pubDate>Sat, 23 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/nix/</guid><description>Nix is a tool for building reproducible environments. It excels in situations where it can have a full view of the world - where every package and piece of configuration is managed by Nix.
One benefit in practice is that it&amp;rsquo;s a really good package manager. Nixpkgs has a ton of packages, and it remains relatively stable even if you follow its &amp;ldquo;unstable&amp;rdquo; branch. Smarter people than me probably know why this is, but I&amp;rsquo;m guessing that the &amp;ldquo;complete view of the world&amp;rdquo; allows Nix to make much stronger guarantees about what will work and what won&amp;rsquo;t.</description><content type="html"><![CDATA[<p><a href="https://nixos.org/">Nix</a> is a tool for building reproducible environments. It excels in situations where it can have a full view of the world - where every package and piece of configuration is managed by Nix.</p>
<p>One benefit in practice is that it&rsquo;s a really good package manager. <a href="https://github.com/NixOS/nixpkgs">Nixpkgs</a> has a ton of packages, and it remains relatively stable even if you follow its &ldquo;unstable&rdquo; branch. Smarter people than me probably know why this is, but I&rsquo;m guessing that the &ldquo;complete view of the world&rdquo; allows Nix to make much stronger guarantees about what will work and what won&rsquo;t.</p>
<p>It&rsquo;s also great at managing user-level configuration. <a href="https://github.com/nix-community/home-manager">Home manager</a> has friendly configuration syntax into a ton of commonly-used programs. Using home manager was my first &ldquo;wtf&rdquo; moment with Nix. I&rsquo;ve been trying to solve the reproducible home problem for <em>so long</em>, and Nix just fixed all my problems.</p>
<p>If you&rsquo;re looking to try out Nix, install it and migrate your home bit-by-bit to home manager! You&rsquo;ll be shocked at how easy it is to set up. If you want to get a little deeper, this <a href="https://nixos-and-flakes.thiscute.world/">book on NixOS and flakes</a> is an amazing resource.</p>
]]></content></item><item><title>Ollama</title><link>https://pawa.lt/braindump/ollama/</link><pubDate>Sat, 23 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/ollama/</guid><description>Ollama is a tool for running LLMs locally, inspired by the DevX of Docker. It has a model hub with the big-name models. You can run models very simply:
$ ollama serve &amp;amp; $ ollama run llama2:latest &amp;gt;&amp;gt;&amp;gt; say whatever you want to say to llama The real magic, though, is in how easy it is to define new model types with different system prompts and parameters:
$ cat Modelfile FROM llama2:latest TEMPERATURE 1 SYSTEM &amp;#34;&amp;#34;&amp;#34;You are an assistant who will write Peyton&amp;#39;s blog posts for him&amp;#34;&amp;#34;&amp;#34; $ ollama create peytonhelper -f Modelfile $ ollama run peytonhelper &amp;gt;&amp;gt;&amp;gt; say something to my helper In the past, I&amp;rsquo;ve used ctransformers for building stuff like this, but it&amp;rsquo;s a bit too low-level.</description><content type="html"><![CDATA[<p><a href="https://ollama.ai/">Ollama</a> is a tool for running LLMs locally, inspired by the DevX of Docker. It has a <a href="https://ollama.ai/library">model hub</a> with the big-name models. You can run models very simply:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ ollama serve &amp;
$ ollama run llama2:latest
&gt;&gt;&gt; say whatever you want to say to llama
</code></pre></div><p>The real magic, though, is in how easy it is to define new model types with different system prompts and parameters:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ cat Modelfile
FROM llama2:latest

TEMPERATURE <span style="color:#ae81ff">1</span>
SYSTEM <span style="color:#e6db74">&#34;&#34;&#34;You are an assistant who will write Peyton&#39;s blog posts for him&#34;&#34;&#34;</span>
$ ollama create peytonhelper -f Modelfile
$ ollama run peytonhelper
&gt;&gt;&gt; say something to my helper
</code></pre></div><p>In the past, I&rsquo;ve used <a href="https://github.com/marella/ctransformers">ctransformers</a> for building stuff like this, but it&rsquo;s a bit too low-level. I don&rsquo;t care to learn the proper prompting setup for all these models.</p>
]]></content></item><item><title>Spacemacs-pilled</title><link>https://pawa.lt/braindump/spacemacs-pilled/</link><pubDate>Sat, 23 Dec 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/braindump/spacemacs-pilled/</guid><description>I don&amp;rsquo;t really care if people use Vim, even if I love it. I care a bit more about if people use &amp;ldquo;vim motions&amp;rdquo;. In my (extremely dogmatic) opinion, vim motions are the only good way to move around code. They make it so easy to move and manipulate text without moving your hands too much.
I don&amp;rsquo;t use Spacemacs, but I feel the same way about it. The way it uses Vim motions to create easy-to-use chords feels so intuitive to me.</description><content type="html"><![CDATA[<p>I don&rsquo;t really care if people use <a href="https://neovim.io/">Vim</a>, even if I love it. I care a bit more about if people use &ldquo;vim motions&rdquo;. In my (extremely dogmatic) opinion, vim motions are the only good way to move around code. They make it so easy to move and manipulate text without <a href="https://my.clevelandclinic.org/health/diseases/17424-repetitive-strain-injury">moving your hands too much</a>.</p>
<p>I don&rsquo;t use <a href="https://www.spacemacs.org/">Spacemacs</a>, but I feel the same way about it. The way it uses Vim motions to create easy-to-use chords feels <em>so</em> intuitive to me. I&rsquo;ve taken the Spacemacs vibe and created my own set of keybindings in the same spirit that work for me.</p>
<p>I use <code>&lt;Space&gt;</code> as my leader with all my actions hanging off of it. Some examples:</p>
<ul>
<li><code>&lt;Space&gt;{h,j,k,l}</code> - move directionally in a pane of split windows</li>
<li><code>&lt;Space&gt;f</code> - symbol search over the project</li>
<li><code>&lt;Space&gt;&lt;Space&gt;</code> - open a pane with all my recently-edited files</li>
</ul>
<p>I&rsquo;ll add more in the future, but for now, you can see some examples in <a href="https://github.com/pawalt/setup/blob/afb377090092f943c71a3f386f4c199b7525185f/homes/common.nix#L181">my VSCode config</a>.</p>
]]></content></item><item><title>Chatting with my recipes</title><link>https://pawa.lt/posts/2023/08/chatting-with-my-recipes/</link><pubDate>Sun, 13 Aug 2023 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2023/08/chatting-with-my-recipes/</guid><description>I&amp;rsquo;ve had a lot of fun playing with LLMs recently. I think they get a lot of coverage in applications where they play the role of an agent, but I&amp;rsquo;ve had pretty mixed results building these kinds of tools. In particular, I&amp;rsquo;ve had trouble getting the LLM to reliably make multiple function calls in serial. While one call may be exactly what I&amp;rsquo;m looking for, there&amp;rsquo;s a decent chance at least one of the calls will be slightly off, breaking the whole flow.</description><content type="html"><![CDATA[<p>I&rsquo;ve had a lot of fun playing with LLMs recently. I think they get a lot of coverage in applications where they play the role of an <a href="https://www.pinecone.io/learn/series/langchain/langchain-agents/">agent</a>, but I&rsquo;ve had pretty mixed results building these kinds of tools. In particular, I&rsquo;ve had trouble getting the LLM to reliably make multiple function calls in serial. While one call may be exactly what I&rsquo;m looking for, there&rsquo;s a decent chance at least one of the calls will be slightly off, breaking the whole flow.</p>
<p>Where I&rsquo;ve had <em>much</em> more reliable results is in using LLMs for extracting structured results from unstructured input and in using them for text manipulation. These both rely less on the LLM being able to &ldquo;think&rdquo; and more on it being able to follow direct instructions.</p>
<h2 id="recipes">Recipes</h2>
<p>In 2021, I made a <a href="/recipes">recipes page</a>, and I seeded it with some recipes. Sadly, I never ended up adding more. While the main reason for this is my laziness, it&rsquo;s also just hard to make good recipes! Formatting ingredients, giving easy-to-follow instructions, and just making the new page is a real pain.</p>
<p>This is a bummer for me since for any of my recipes, I can easily articulate in words how to make it, but I find writing those instructions down very difficult. Recently, I&rsquo;ve turned to voice recognition and LLMs for assistance on this.</p>
<p>At a high level, I want to enable the following workflow:</p>
<ul>
<li><input disabled="" type="checkbox"> Dictate a recipe using my voice</li>
<li><input disabled="" type="checkbox"> Have an LLM generate a nicely-formatted recipe</li>
<li><input disabled="" type="checkbox"> Put that recipe on my website</li>
<li><input disabled="" type="checkbox"> Dictate any modifications I want to make and have them reflected in the recipe.</li>
</ul>
<h2 id="voice-transcription">Voice Transcription</h2>
<p>The first step of all of this is to go from voice -&gt; text. I don&rsquo;t have much familiarity with recording in Python, so I mostly asked ChatGPT to write the code and told it what was broken. <a href="https://github.com/pawalt/beyn/blob/05c7b6c6468b0462d8a0edd41af62f575492dc23/listen.py">The code for this</a> isn&rsquo;t all that interesting, but it provides a nice interface that I can use to record audio:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python"><span style="color:#75715e"># This will record audio to the file listen.RECORDING_FILE</span>
listen<span style="color:#f92672">.</span>record_audio()
</code></pre></div><p>To go from this recording file to text, I used a <a href="https://github.com/stlukey/whispercpp.py">python port</a> of the incredible <a href="https://github.com/ggerganov/whisper.cpp">whisper.cpp</a> project. I am consistently stunned at how well whisper.cpp works and how quickly it runs. It outperforms the in-built MacOS dictation by such a wide margin that you&rsquo;ll never want to use Siri again.</p>
<p>With whisper set up, I now have a super simple function I can use to get voice input:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python"><span style="color:#66d9ef">def</span> <span style="color:#a6e22e">get_voice_input</span>(whisp: Whisper) <span style="color:#f92672">-&gt;</span> str:
    listen<span style="color:#f92672">.</span>record_audio()
    <span style="color:#66d9ef">return</span> whisp<span style="color:#f92672">.</span>transcribe_from_file(listen<span style="color:#f92672">.</span>RECORDING_FILE)
</code></pre></div><p>Here&rsquo;s where we are so far:</p>
<ul>
<li><input checked="" disabled="" type="checkbox"> Dictate a recipe using my voice</li>
<li><input disabled="" type="checkbox"> Have an LLM generate a nicely-formatted recipe</li>
<li><input disabled="" type="checkbox"> Put that recipe on my website</li>
<li><input disabled="" type="checkbox"> Dictate any modifications I want to make and have them reflected in the recipe.</li>
</ul>
<h2 id="function-calls">Function calls</h2>
<p>I alluded to function calls before, but it&rsquo;s worth taking some time to talk about what they really are. In June, <a href="https://openai.com/blog/function-calling-and-other-api-updates">OpenAI announced an extension</a> to their API to allow GPT chat completions to call functions instead of just responding with text. You simply provide a JSON schema describing the available functions, and the LLM will make a decision on whether to call a function and what arguments to call it with.</p>
<p>This has primarily been billed as a tool to build agents. For example, if you want to build an agent that can tell you the weather, you can give it a &ldquo;get_current_weather&rdquo; function that&rsquo;ll make a call out to some weather API. Then, when you ask it &ldquo;what&rsquo;s the weather&rdquo;, it&rsquo;ll make the function call and respond to your question in plain english using the results of the call.</p>
<p>The powerful thing here is that the function calling is <em>extremely reliable</em> in only ever making calls that conform to the JSON schema. I&rsquo;ve had very few issues, and I was able to fix any issues by prompting the LLM more clearly.</p>
<h3 id="autogenerated-schemas">Autogenerated schemas</h3>
<p>JSON schemas are pretty arduous to type out, in particular because you have to define your schema once in code and once in the JSON schema you provide to the LLM. To alleviate this pain, <a href="https://github.com/jxnl/openai_function_call">there&rsquo;s a library</a> that generates OpenAI-compatible JSON schemas straight from Pydantic objects. This allows you to define type-safe schemas that LLMs can understand &ldquo;for free&rdquo;.</p>
<p>From the docs, here&rsquo;s an example of using the library to extra user information from input:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python"><span style="color:#66d9ef">class</span> <span style="color:#a6e22e">UserDetails</span>(OpenAISchema):
    <span style="color:#e6db74">&#34;&#34;&#34;User Details&#34;&#34;&#34;</span>
    name: str <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;User&#39;s name&#34;</span>)
    age: int <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;User&#39;s age&#34;</span>)

completion <span style="color:#f92672">=</span> openai<span style="color:#f92672">.</span>ChatCompletion<span style="color:#f92672">.</span>create(
    model<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;gpt-3.5-turbo-0613&#34;</span>,
    functions<span style="color:#f92672">=</span>[UserDetails<span style="color:#f92672">.</span>openai_schema],
    messages<span style="color:#f92672">=</span>[
        {<span style="color:#e6db74">&#34;role&#34;</span>: <span style="color:#e6db74">&#34;system&#34;</span>, <span style="color:#e6db74">&#34;content&#34;</span>: <span style="color:#e6db74">&#34;I&#39;m going to ask for user details. Use UserDetails to parse this data.&#34;</span>},
        {<span style="color:#e6db74">&#34;role&#34;</span>: <span style="color:#e6db74">&#34;user&#34;</span>, <span style="color:#e6db74">&#34;content&#34;</span>: <span style="color:#e6db74">&#34;My name is John Doe and I&#39;m 30 years old.&#34;</span>},
    ],
)

user_details <span style="color:#f92672">=</span> UserDetails<span style="color:#f92672">.</span>from_response(completion)
<span style="color:#66d9ef">print</span>(user_details)  <span style="color:#75715e"># UserDetails(name=&#34;John Doe&#34;, age=30)</span>
</code></pre></div><p>You might be able to see where I&rsquo;m going with this :)</p>
<h2 id="recipe-extraction">Recipe extraction</h2>
<p>In order to render a recipe onto a page, I need to know its components. I thought the following were pretty typical attributes of a recipe:</p>
<ul>
<li>Description</li>
<li>Steps</li>
<li>Ingredients</li>
<li>Total Time</li>
<li>Active Time</li>
</ul>
<p>Using the function call library, I wrote up a <a href="https://github.com/pawalt/beyn/blob/05c7b6c6468b0462d8a0edd41af62f575492dc23/elems.py">schema for recipes</a>:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python"><span style="color:#66d9ef">class</span> <span style="color:#a6e22e">RecipeDetails</span>(OpenAISchema):
    <span style="color:#e6db74">&#34;&#34;&#34;RecipeDetails are the details describing a recipe. Details are precise and
</span><span style="color:#e6db74">reflect the input from the user.&#34;&#34;&#34;</span>
    description: str <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;casual description of every recipe step in approximately 200 characters&#34;</span>)
    steps: List[str] <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;&#34;&#34;specific list of all steps in the recipe, in order.
</span><span style="color:#e6db74">If multiple steps can be combined into one, they will.&#34;&#34;&#34;</span>)
    recipe_title: str <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;title of the recipe in 20 characters or less&#34;</span>)
    ingredients: List[Ingredient] <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;list of all ingredients in the recipe&#34;</span>)
    total_time: int <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;total time in minutes required for this recipe&#34;</span>)
    active_time: int <span style="color:#f92672">=</span> Field(<span style="color:#f92672">...</span>, description<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;total active (non-waiting) cooking time in minutes required for this recipe&#34;</span>)
</code></pre></div><p>I then wrote a generic function to extract any schema from unstructured text and called it on the transcript:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python">transcription <span style="color:#f92672">=</span> get_voice_input(w)
recipe <span style="color:#f92672">=</span> llm<span style="color:#f92672">.</span>get_shema_obj(<span style="color:#e6db74">&#34;RecipeDetails&#34;</span>, elems<span style="color:#f92672">.</span>RecipeDetails, transcription)
</code></pre></div><p>Now, <code>recipe</code> is a structured object that I can pull all of my recipe information out of.</p>
<ul>
<li><input checked="" disabled="" type="checkbox"> Dictate a recipe using my voice</li>
<li><input checked="" disabled="" type="checkbox"> Have an LLM generate a nicely-formatted recipe</li>
<li><input disabled="" type="checkbox"> Put that recipe on my website</li>
<li><input disabled="" type="checkbox"> Dictate any modifications I want to make and have them reflected in the recipe.</li>
</ul>
<h2 id="rendering-the-recipe">Rendering the recipe</h2>
<p>I already use Hugo to host my blog, so I won&rsquo;t fix what ain&rsquo;t broke. First, when I create a new recipe, I make a small markdown file inside of my <code>content/recipes</code> directory with some metadata and a reference to the JSON-formatted recipe info:</p>
<pre><code class="language-hugo" data-lang="hugo">+++
title = &quot;The Good Lemonade&quot;
description = &quot;Refreshing and sweet lemonade like they served at the country fair&quot;
date = &quot;2023-08-06&quot;
total_time = &quot;1440&quot;
active_time = &quot;20&quot;
+++

{{&lt; ai_recipe url=&quot;data/recipes/country_fair_lemonade.json&quot; &gt;}}
</code></pre><p>This <code>ai_recipe</code> shortcode is a <a href="https://github.com/pawalt/personal-site/blob/d69ef25b15a1cc95d43cdf7aed2715f4e62bc407/layouts/shortcodes/ai_recipe.html">pretty simple template</a> that displays all the elements of the JSON recipe. <a href="https://github.com/pawalt/personal-site/blob/d69ef25b15a1cc95d43cdf7aed2715f4e62bc407/data/recipes/country_fair_lemonade.json">The JSON</a> is just the <code>RecipeDetails</code> class serialized.</p>
<p>And that&rsquo;s it for the first pass! I can now speak into my mic and get a nicely-formatted recipe on my website!</p>
<iframe width="560" height="315" src="https://www.youtube.com/embed/cjF5rj-5NuA" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen></iframe>
<ul>
<li><input checked="" disabled="" type="checkbox"> Dictate a recipe using my voice</li>
<li><input checked="" disabled="" type="checkbox"> Have an LLM generate a nicely-formatted recipe</li>
<li><input checked="" disabled="" type="checkbox"> Put that recipe on my website</li>
<li><input disabled="" type="checkbox"> Dictate any modifications I want to make and have them reflected in the recipe.</li>
</ul>
<h2 id="modifications">Modifications</h2>
<p>These recipes are never perfect in the first iteration. The ingredients are typically overspecified, and the steps definitely suffer from miscommunications.</p>
<p>To solve this, I built a new mode into my REPL for performing modifications. I just ask the user for a link to the recipe they&rsquo;d like to modify, and I pull the unique identifier out of the URL. I then read in the JSON for that recipe so I can kick off modification:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python">url <span style="color:#f92672">=</span> nice_input(<span style="color:#e6db74">&#34;give the url for the recipe: &#34;</span>)
<span style="color:#75715e"># Regular expression pattern</span>
pattern <span style="color:#f92672">=</span> <span style="color:#e6db74">r</span><span style="color:#e6db74">&#39;https?://[^/]+/recipes/ai_([^/]+)/?&#39;</span>

<span style="color:#75715e"># Use re.search to find matches</span>
match <span style="color:#f92672">=</span> re<span style="color:#f92672">.</span>search(pattern, url)

<span style="color:#66d9ef">if</span> match:
	recipe_name <span style="color:#f92672">=</span> match<span style="color:#f92672">.</span>group(<span style="color:#ae81ff">1</span>)

	json_path <span style="color:#f92672">=</span> os<span style="color:#f92672">.</span>path<span style="color:#f92672">.</span>join(presentation<span style="color:#f92672">.</span>hugo_base_dir, <span style="color:#e6db74">&#34;data&#34;</span>, <span style="color:#e6db74">&#34;recipes&#34;</span>, f<span style="color:#e6db74">&#34;{recipe_name}.json&#34;</span>)
	<span style="color:#66d9ef">with</span> open(json_path, <span style="color:#e6db74">&#39;r&#39;</span>) <span style="color:#66d9ef">as</span> f:
		file_contents <span style="color:#f92672">=</span> f<span style="color:#f92672">.</span>read()
	recipe_dict <span style="color:#f92672">=</span> json<span style="color:#f92672">.</span>loads(file_contents)
	recipe <span style="color:#f92672">=</span> parse_obj_as(elems<span style="color:#f92672">.</span>RecipeDetails, recipe_dict)
</code></pre></div><h3 id="modification-repl">Modification REPL</h3>
<p>In order to make modifications relative to the current state, I need to show the LLM what the current state of the file is. I don&rsquo;t want to pass in a full HTML page, so I made a text analog of my webpage using jinja:</p>
<pre><code class="language-jinja2" data-lang="jinja2">Recipe Title: {{ recipe.recipe_title }}

Total time: {{ recipe.total_time }} minutes
Active time: {{ recipe.active_time }} minutes

Recipe Description: {{ recipe.description }}

Recipe Ingredients:
{% for item in recipe.ingredients -%}
- {{ item.quantity }} {{ item.unit }} {{ item.name }}
{% endfor %}
Recipe steps:
{% for step in recipe.steps -%}
- {{ step }}
{% endfor %}
</code></pre><p>I then ask the user for voice input and instruct the LLM to make modifications to the recipe according to the user&rsquo;s instructions:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python">modifications <span style="color:#f92672">=</span> get_voice_input(w)
new_recipe <span style="color:#f92672">=</span> llm<span style="color:#f92672">.</span>make_modifications(modifications, recipe)
<span style="color:#75715e"># use the same parsing as before to turn unstructured recipe into pydantic class</span>
recipe <span style="color:#f92672">=</span> parse_recipe(new_recipe)
</code></pre></div><p>I can then write out the new recipe, and the user can look in their live Hugo view to see if it looks good:</p>
<iframe width="560" height="315" src="https://www.youtube.com/embed/kvX0AklCPRw" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen></iframe>
<p>There are some cases where the LLM will make modifications too aggressively or it&rsquo;s just easier for me to make these changes manually. For these cases, <a href="https://github.com/pawalt/beyn/blob/05c7b6c6468b0462d8a0edd41af62f575492dc23/main.py#L24">I open up a vim buffer</a> with the recipe and allow the user to make whatever changes they wish. I use the LLM to parse the required info out the same as before.</p>
<ul>
<li><input checked="" disabled="" type="checkbox"> Dictate a recipe using my voice</li>
<li><input checked="" disabled="" type="checkbox"> Have an LLM generate a nicely-formatted recipe</li>
<li><input checked="" disabled="" type="checkbox"> Put that recipe on my website</li>
<li><input checked="" disabled="" type="checkbox"> Dictate any modifications I want to make and have them reflected in the recipe.</li>
</ul>
<h2 id="conclusion">Conclusion</h2>
<p>This tool was a ton of fun to build and taught me to think more precisely about how to use LLMs. I really believe that the current assistant/agent-focused discourse is pretty misguided. While assistant applications may come into the mainstream, I think that in the near term, LLMs will be useful as tools to get jobs done rather than agents to replace humans.</p>
<h3 id="local-llm-epilogue">Local LLM epilogue</h3>
<p>I&rsquo;ve been playing around with local LLMs a lot as well. They are not even close to GPT quality, but they&rsquo;re getting better every day. Hopefully Llama 2&rsquo;s commercial license will open the floodgates.</p>
<p>With <a href="https://github.com/ggerganov/llama.cpp/pull/1773">grammar based sampling</a> being merged into llama.cpp, I think we&rsquo;re a JSON-trained model away from being able implement function calling locally. <a href="https://github.com/go-skynet/LocalAI/issues/588">Some have already taken a stab at it</a>, and I have high hopes for restricted sampling being the key to eeking out better performance from local models.</p>
<p>At some point in the future, I&rsquo;d like to rip out the OpenAI calls here and replace them with a local model. That would be sweet.</p>
]]></content></item><item><title>Golinks at home</title><link>https://pawa.lt/posts/2022/07/golinks-at-home/</link><pubDate>Tue, 12 Jul 2022 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2022/07/golinks-at-home/</guid><description>I recently moved into an apartment which means I get to have something I&amp;rsquo;ve missed for the last four years: a homelab! After hearing about the homelab, my roommate asked me &amp;ldquo;Would it be possible to get network-wide golinks?&amp;rdquo; After some thinking and looking at a similar project I worked on, I decided to build them!
How Golinks Work The goal of a golink is to allow the user to input a short URL like go/gh/ and be redirected to a longer url such as https://github.</description><content type="html"><![CDATA[<p>I recently moved into an apartment which means I get to have something I&rsquo;ve missed for the last four years: a homelab! After hearing about the homelab, my roommate asked me &ldquo;Would it be possible to get network-wide golinks?&rdquo; After some thinking and looking at a <a href="https://github.com/pawalt/shortname">similar project I worked on</a>, I decided to build them!</p>
<h2 id="how-golinks-work">How Golinks Work</h2>
<p>The goal of a golink is to allow the user to input a short URL like <code>go/gh/</code> and be redirected to a longer url such as <code>https://github.com/</code>. These should also be path-preserving, so <code>go/gh/pawalt</code> should go to <code>https://github.com/pawalt</code>.</p>
<p>This is primarily achieved by reconfiguring DNS with a special mapping for <code>go</code> and serving a redirect whenever we see one of our redirect paths.</p>
<h3 id="dns-lookup">DNS lookup</h3>
<p>When the browser (or any HTTP client) looks up a URL like <code>go/gh</code>, it first has to know what IP to send the request to. To do this, it looks up the host of the request (in this case <code>go</code>) using a DNS server. We need to get some access to our DNS server so that we can set the <code>go</code> entry to our IP.</p>
<p>Additionally, if we want to serve redirects on hosts other than <code>go</code> (<code>gh -&gt; github.com</code> for example), we&rsquo;ll need to update DNS with each host that we want to redirect from.</p>
<h3 id="serving-the-redirect">Serving the redirect</h3>
<p>Now that a request for <code>go/</code> has made it to our application, we need a way to serve a redirect to our destination URL. In order to serve the redirect, we&rsquo;ll run a web server that accepts request on all routes. When it gets a request, it looks up the host and path in its list of redirects. If it finds a matching redirect rule, it&rsquo;ll serve an HTTP redirect (usually 302) to the destination URL.</p>
<p>The combination of DNS + HTTP redirect means requests for any source URL will get automatically redirected to their destination URLs.</p>
<h3 id="editing-the-redirect-list">Editing the redirect list</h3>
<p>With the mechanics in place, we need a way to edit the mappings from source to destination URL. For these, we can simply run a web server that retrieves a list of the current mappings and displays them, allowing the user to edit the list. When the user saves a new set of mappings, it must be stored ✨ somewhere ✨, and DNS must be reconfigured for any new hosts that are added.</p>
<p>With each component, we get a workflow that looks like:</p>
<ol>
<li>User adds a new redirect in web editor</li>
<li>New redirect is saved, and DNS is reconfigured with any new hosts</li>
<li>User can make a request and be redirected
<ol>
<li>DNS points the host to the right IP</li>
<li>The webserver serves a redirect to the desired URL</li>
</ol>
</li>
</ol>
<h2 id="implementation">Implementation</h2>
<p>To implement this, I&rsquo;ve got two devices I&rsquo;m talking to - my router for the DNS and the computer running the redirect webserver. In this case, I&rsquo;m using a <a href="https://fly.io/">fly.io</a> VM to host the redirect webserver. Each of these devices might be on completely separate networks with no way to talk to each other.</p>
<figure>
    <img src="/img/golinks_1.png"
         alt="Network diagram without connections"/> <figcaption>
            <p>Network diagram without connections</p>
        </figcaption>
</figure>

<h3 id="tailscale">Tailscale</h3>
<p>To connect my devices, I&rsquo;m using <a href="https://tailscale.com/">Tailscale</a>. You can think of Tailscale as a zero-config peer-to-peer VPN solution, although it&rsquo;s really more than that (words on this later). Tailscale gives each of my devices an IP in the 100.x.y.z range and connects them using a WireGuard tunnel. This means that even on different networks, my laptop can directly access both devices on consistent IPs.</p>
<figure>
    <img src="/img/golinks_2.png"
         alt="Network diagram with Tailscale-provisioned 100.x.y.z IPs"/> <figcaption>
            <p>Network diagram with Tailscale-provisioned 100.x.y.z IPs</p>
        </figcaption>
</figure>

<p>With a little extra work, we can even get this connected to the entire home network, meaning non-Tailscale devices can still use golinks. To do this, I designate my router as a <a href="https://tailscale.com/kb/1019/subnets/">subnet router</a>, meaning that Tailscale knows to route traffic for my home network through it. In this case, my home network is <code>172.27.0.0/16</code>, so any requests for that range will go through my router, <code>100.2.3.4</code>.</p>
<figure>
    <img src="/img/golinks_2.5.png"
         alt="Home network connected to tailnet via router"/> <figcaption>
            <p>Home network connected to tailnet via router</p>
        </figcaption>
</figure>

<h3 id="host-file">Host file</h3>
<p>First, we need to get DNS overriding working - requests for URLs like <code>go/</code> should be sent to our VM in Fly. We can achieve this by adding a hosts file to our router&rsquo;s DNS settings. I&rsquo;m using <a href="https://openwrt.org/start">OpenWRT</a> on my router which uses <a href="https://thekelleys.org.uk/dnsmasq/doc.html">Dnsmasq</a> for DHCP and DNS. Just as <code>/etc/hosts</code> lets us define custom <code>host -&gt; IP</code> mappings on our computers, we can add a hosts file to Dnsmasq for our own custom mappings.</p>
<p>The hosts file will use the Fly VM&rsquo;s tailscale IP as the IP to connect to, and it&rsquo;ll store the mappings in comments. This information is our entire state which means we can just use this file instead of any databases. Here&rsquo;s how a table of redirects would be formatted down in our hosts file:</p>
<table>
<thead>
<tr>
<th><strong>Source</strong></th>
<th><strong>Destination</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>gh</td>
<td><a href="https://github.com">https://github.com</a></td>
</tr>
<tr>
<td>gh/pl</td>
<td><a href="https://github.com/pennlabs">https://github.com/pennlabs</a></td>
</tr>
<tr>
<td>gh/pw</td>
<td><a href="https://github.com/pawalt">https://github.com/pawalt</a></td>
</tr>
<tr>
<td>go/gh</td>
<td><a href="https://github.com">https://github.com</a></td>
</tr>
<tr>
<td>go/mon</td>
<td><a href="https://www.youtube.com/watch?v=b2F-DItXtZs">https://www.youtube.com/watch?v=b2F-DItXtZs</a></td>
</tr>
</tbody>
</table>
<pre><code>root@OpenWrt:~# cat /etc/hosts.d/golinks.hosts
# this hosts file property of me!!! dont touch


100.79.24.134 gh # gh -&gt; https://github.com
100.79.24.134 gh # gh/pl -&gt; https://github.com/pennlabs
100.79.24.134 gh # gh/pw -&gt; https://github.com/pawalt
100.79.24.134 go # go/gh -&gt; https://github.com
100.79.24.134 go # go/mon -&gt; https://www.youtube.com/watch?v=b2F-DItXtZs
</code></pre><h3 id="serving-the-redirect-1">Serving the redirect</h3>
<p>Serving the redirect is fairly straightforward - when a user makes a web request to the Fly VM, it will look up the link in its redirect map and serve a redirect (302 Found) to the redirect destination. If it can&rsquo;t find a matching redirect, it&rsquo;ll serve a 404.</p>
<h3 id="editing-the-redirect-list-1">Editing the redirect list</h3>
<p>To view and edit the redirects, the same webserver that does the redirecting serves a UI on <code>go/_/hosts</code>. I use a weird path (<code>/_/hosts</code>) so there&rsquo;s a low chance of collision with any URL I&rsquo;d actually want to go to via a golink.</p>
<p><img src="/img/golinks_3.png" alt="UI to edit list of redirects">
When the user hits <code>Submit</code>, the list of redirects is POSTed to <code>/_/hosts</code>. The list is then persisted in the router&rsquo;s golinks host file (mine is <code>/etc/hosts.d/golinks.hosts</code>).</p>
<h4 id="mapping-persistence">Mapping persistence</h4>
<p>As previously mentioned, the information on the mappings is stored entirely in the router&rsquo;s hosts file so as to not have any local state to deal with. In an ideal world, I could just use SFTP to read and write the mapping data in the hosts file, but this normally requires me to mint an ssh key and distribute it to both Fly and my router.</p>
<p>Luckily, <a href="https://tailscale.com/blog/tailscale-ssh/">Tailscale SSH</a> obviates the need for any of that! By running Tailscale SSH on my router<sup id="fnref:1"><a href="#fn:1" class="footnote-ref" role="doc-noteref">1</a></sup>, I can SSH into it via any device I&rsquo;m logged into. This means that my Fly VM can SSH into the router without providing any credentials - simply by being on the tailnet, it has the authorization to SSH<sup id="fnref:2"><a href="#fn:2" class="footnote-ref" role="doc-noteref">2</a></sup>.</p>
<p><strong>Reading</strong></p>
<p>To read in a mapping, the application will SFTP into the router and read the hosts file (default <code>/etc/hosts.d/golinks.hosts</code>). It&rsquo;ll then parse out the comments and update its stored mapping to reflect the mappings in the comments.</p>
<p>This list is refreshed every 5 minutes in case of a out-of-band edit to the file.</p>
<p><strong>Writing</strong></p>
<p>Before writing out the hosts file, I have to answer a question: what IP should be advertised to the clients i.e. what IP does the Fly VM have that is routable from my laptop? It&rsquo;s the Tailscale IP! To get the Tailscale IP, I dial a UDP connection to <code>100.100.100.100</code>, one of Tailscale&rsquo;s managed IPs. Since it&rsquo;s Tailscale-managed, the source IP on this connection is my Tailscale IP.</p>
<p>To render out the hosts file, I have a very simple template:</p>
<pre><code># this hosts file property of me!!! dont touch  
  
{{ range .lines }}  
{{ .ip }} {{ .host }} # {{ .comment }}  
{{- end }}
</code></pre><p>I render this template and SFTP it over to the router to save the new configuration, but there&rsquo;s one last trick. Dnsmasq won&rsquo;t immediately see the new host file, so it might be a few minutes before I can use my link. However, if Dnsmasq receives a <code>SIGHUP</code>, it will reload all its config - including the hosts files. To send the <code>SIGHUP</code>, I use a quick command: <code>kill -HUP $(ps | grep dnsmasq | grep -v grep | cut -d &quot; &quot; -f1)</code>. There&rsquo;s definitely a better way to do this with <code>syscall.Kill</code>, but that&rsquo;s a task for later.</p>
<p>Now that we have persistence, our end to end process looks like:</p>
<ol>
<li>User adds a new redirect at <code>go/_/hosts</code></li>
<li>New redirect is saved via sftp to <code>/etc/hosts.d/golinks.hosts</code>, reconfiguring Dnsmasq</li>
<li>User can make a request and be redirected
<ol>
<li>DNS points the host to the Fly VM</li>
<li>The webserver serves a redirect to the desired URL</li>
</ol>
</li>
</ol>
<video style="max-width: 100%;" controls>
  <source src="/img/golinks_demo.mp4" type="video/mp4">
Your browser does not support the video tag.
</video>
<h2 id="conclusion">Conclusion</h2>
<p>In a small bit of code, we&rsquo;ve managed to get network-wide golinks and keep them on the go! I&rsquo;ve got Tailscale installed on all my devices, so this solution works wherever I am on whatever device I&rsquo;m using. This new SSH feature has changed the way I see the product - from a VPN solution to a network mesh with trust built in. I hope to see more features like this in the future!</p>
<p>If you want to check out any of the code for this project, head over to <a href="https://github.com/pawalt/homelab/tree/main/golinks">my homelab repo</a>. Just be warned: the code is definitely homelab-quality.</p>
<section class="footnotes" role="doc-endnotes">
<hr>
<ol>
<li id="fn:1" role="doc-endnote">
<p>Unfortunately, the <a href="https://openwrt.org/packages/pkgdata/tailscale">tailscale OpenWRT package</a> is pretty out-of-date as of this post. To get Tailscale SSH, I SFTPed the <a href="https://pkgs.tailscale.com/stable/#static">Tailscale ARM statically-linked binary</a> to my router. <a href="#fnref:1" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
<li id="fn:2" role="doc-endnote">
<p>I&rsquo;m currently on a single-user plan, so my rules are pretty relaxed. Once Tailscale gets <a href="https://tailscale.com/kb/1064/invite-team-members/#how-can-i-invite-someone-if-i-signed-up-with-a-gmail-address-or-a-github-personal-account">better account sharing</a>, I&rsquo;ll do some ACL work and lock down SSH access to only my personal devices. <a href="#fnref:2" class="footnote-backref" role="doc-backlink">&#x21a9;&#xfe0e;</a></p>
</li>
</ol>
</section>
]]></content></item><item><title>Teenage Anxiety</title><link>https://pawa.lt/posts/2022/05/teenage-anxiety/</link><pubDate>Wed, 25 May 2022 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2022/05/teenage-anxiety/</guid><description>Right before my junior year of high school, I hit a very lucky break and got the opportunity to work at a datacenter part-time during school. Considering that I worked a horrendous tutoring job at the time, this was the opportunity of a lifetime for me.
The Work As soon as I started this job, I immediately loved the work. I latched onto networking and sysadmin work in a way that I had never latched onto any schoolwork in my life.</description><content type="html"><![CDATA[<p>Right before my junior year of high school, I hit a very lucky break and got the opportunity to work at <a href="https://richweb.com/">a datacenter</a> part-time during school. Considering that I worked a <strong>horrendous</strong> tutoring job at the time, this was the opportunity of a lifetime for me.</p>
<h2 id="the-work">The Work</h2>
<p>As soon as I started this job, I immediately loved the work. I latched onto networking and sysadmin work in a way that I had never latched onto any schoolwork in my life. Work provided me constant challenges that I could bash my head against and eventually overcome. When I conquered one challenge, the next was sitting there, waiting for me to run it over.</p>
<p>Additionally, I received praise like I had never before in my life. In school I was a high performer, but I was rarely top of my class. Even when I was, it never felt gratifying. Grades have never been a motivating factor for me. They&rsquo;ve always felt secondary to the main goal - learning the content. At work however, I got to work on hard problems, and when I solved them, the people around me were impressed and praised me for my work.</p>
<h2 id="the-anxiety">The Anxiety</h2>
<p>This praise was not as simple as &ldquo;great job.&rdquo; Most frequently it was along the lines of &ldquo;great job - aren&rsquo;t you 16??&rdquo; At the beginning, this praise was strictly a positive influence on me, but as I grew, it haunted me. What if I could no longer achieve that level of praise? What if as I grew older, I lost my &ldquo;wonder kid magic&rdquo;?</p>
<p><strong>As I grew, the happiness I got from that praise turned to anxiety.</strong> Instead of seeking to do the best job for the company and our clients, I sought to gain the maximum amount of praise possible. Instead of admitting where I had doubts about my solutions, I glossed over my doubts and presented my work as flawless.</p>
<h2 id="the-issues">The Issues</h2>
<p>Perhaps one of my best examples was when I was working on installing <a href="https://www.dpdk.org/">kernel bypass routing software</a> on one of our routers. I worked on installing this software all morning and never got it working 100% right. Before I went to lunch, though, I told my boss we were all good and that I hadn&rsquo;t caused any issues. I came back from lunch to my boss asking &ldquo;Hey Peyton, any idea why half our sites went down an hour ago?&rdquo; I had glossed over issues before this, but this is one of the only times it actually bit me in the ass.</p>
<p>I was anxious not only about my performance but also about getting older! So much of my praise at the time was contextualized as &ldquo;I can&rsquo;t believe you did X as a Y year old!&rdquo; My greatest fear was that one day I would wake up and no longer be that wonder child. I feared growing older and just being some random dude, indistinguishable from the millions of other engineers out there.</p>
<h2 id="growth">Growth</h2>
<p>As I&rsquo;ve grown, I&rsquo;ve realized that fear of age was unfounded. The older I get, the more I realize that being good at what I do is enough. I don&rsquo;t need to be some special kid; I just need to be good at what I do and to love what I do.</p>
<p>I&rsquo;ve also realized that sweeping my sketchy implementation details under the rug actually makes me worse at my job! These details compound, and before long, the collective burden of small issues becomes unbearable. Building solid systems that stand the test of time is ultimately more fulfilling, even if I have to admit that I don&rsquo;t always know what to do.</p>
<p>Some of these realizations came after working on <a href="https://github.com/cockroachdb/cockroach">CockroachDB</a> this summer. When working on an open-source codebase, there is nowhere to hide. All of my work was open to the public, and as such, it was scrutinized thoroughly. I frequently thought to myself &ldquo;eh this isn&rsquo;t the cleanest way to do things but it&rsquo;ll work.&rdquo; Literally every time I thought that, the &ldquo;clean way&rdquo; was brought up in code review.</p>
<p>Additionally, the <a href="https://github.com/cockroachdb/cockroach/pull/67969">RFC process</a> forced me to think out <em>every</em> detail of how my solution would work before implementing it. This was an extremely rewarding experience as instead of sweeping details under the rug as I might&rsquo;ve done before, I got the opportunity to receive feedback on all the minutiae of my design from people far more knowledgeable than me.</p>
<p>I look forward to many more years of building robust systems and being honest as I build them.</p>
]]></content></item><item><title>Path@Penn Downtime Post-Mortem</title><link>https://pawa.lt/posts/2022/05/pathpenn-downtime-post-mortem/</link><pubDate>Wed, 04 May 2022 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2022/05/pathpenn-downtime-post-mortem/</guid><description>NOTE: This post was cross-posted from the Penn Labs blog! Check it out over there as well and stay around to see some of the cool work the team does: https://pennlabs.org/blog/spot-postmortem
Over the past month, our infrastructure has had some intermittent issues, so we wanted to drop a quick postmortem describing what the problem was and how we fixed it so we can solve these problems more quickly in the future!</description><content type="html"><![CDATA[<blockquote>
<p>NOTE: This post was cross-posted from the Penn Labs blog! Check it out over there as well and stay around to see some of the cool work the team does:
<a href="https://pennlabs.org/blog/spot-postmortem">https://pennlabs.org/blog/spot-postmortem</a></p>
</blockquote>
<p>Over the past month, our infrastructure has had some intermittent issues, so we wanted to drop a quick postmortem describing what the problem was and how we fixed it so we can solve these problems more quickly in the future!</p>
<h2 id="terminology">Terminology</h2>
<p>Before starting, we should define some terminology:</p>
<ul>
<li><strong>Pod</strong> - one instance of a running application (ex. Penn Courses Backend)</li>
<li><strong>Node</strong> - a physical machine that our pods run on</li>
<li><strong>Cluster</strong> - a collection of nodes that our pods run on</li>
<li><strong>Kubernetes</strong> - a piece of software built to move pods among nodes such that nodes don&rsquo;t get overwhelmed (use too much CPU/RAM)</li>
</ul>
<h2 id="background">Background</h2>
<p>For our cluster, we use <a href="https://aws.amazon.com/ec2/spot/">AWS Spot Nodes</a>. These are machines that may, at any time, be shut down by AWS for use in other places. In exchange for this, we pay a lower price for these nodes. Nodes going down is not a problem for us since we&rsquo;re running in Kubernetes which will automatically move our applications to our other nodes if one goes down.</p>
<p>When a spot node goes down, all applications (I&rsquo;ll use the word pods interchangeably) are moved to the other nodes. A new node is then brought up and the pods (hopefully) rebalance between the available nodes.</p>
<p>Additionally, a few weeks ago, we switched from many smaller nodes to 3 higher-resource nodes. This gave us some cost savings as fewer high-resource nodes are cheaper than many low-resource nodes.</p>
<h2 id="symptoms">Symptoms</h2>
<p>Symptoms first appeared around the time of the Path@Penn launch.</p>
<p>We started getting reports of applications failing with no text in the browser other than <code>Service Unavailable</code>. This allowed us to immediately know that this was an infrastructure-wide issue for a few reasons:</p>
<ul>
<li>If multiple products are failing in the same way, then there cannot be a single-application code change that caused this. It must be that some part of the underlying infrastructure is broken.</li>
<li><code>Service Unavailable</code> is an error that typically only our load balancer, <a href="https://traefik.io/">Traefik</a> will throw. It is almost never sent by actual application code.</li>
</ul>
<h2 id="initial-debugging">Initial Debugging</h2>
<p>Upon looking at the cluster with <code>kubectl</code>, we saw that many pods were stuck in the <code>ContainerCreating</code> state. This means that a pod has been assigned to a node, but for some reason, the node cannot start the pod up.</p>
<p>To debug this further, we looked at Datadog, our monitoring system. Datadog showed between 1 and 2 nodes at 100% CPU usage with the third sitting almost idle at ~10% usage.</p>
<p>At this point, the question becomes: <strong>If Kubernetes is balancing pods, why are some nodes at such higher resource usage than the others?</strong></p>
<h2 id="deduction">Deduction</h2>
<p>After further investigation, it looks like the following was happening:</p>
<ol>
<li>A spot node gets scheduled to terminate</li>
<li>All its pods are drained to the other two nodes</li>
<li>The remaining two nodes don&rsquo;t have enough available resources to host all our pods and spike to 100% CPU usage</li>
<li>100% CPU usage renders the two nodes unusable, and they stop responding to both Kubernetes and Datadog.</li>
<li>Even when the third node comes back online, pods cannot be rebalanced to it because the other two nodes are unreachable due to load.</li>
</ol>
<p>The underlying cause of this issue is that while 3 nodes is enough for our cluster, 2 is not. When a spot node termination happens, we don&rsquo;t have enough resources to handle the load, and it creates a cascading failure throughout our infrastructure.</p>
<h2 id="solution">Solution</h2>
<p>To solve this issue, we&rsquo;ve scaled our cluster up to 5 nodes. This should mean that even if we lose a node due to spot lifecycle, we&rsquo;ve got plenty of compute to handle things.</p>
<h2 id="why-did-this-just-happen-now">Why did this just happen now?</h2>
<p>This is the main question we&rsquo;ve been asking ourselves. Intuitively, it seems like we should have run into these issues as soon as we switched to only using 3 nodes.</p>
<p>However, these issues are not just a symptom of having 3 nodes. <strong>They&rsquo;re a symptom of having 3 nodes while under high load</strong>. The Path@Penn launch was the first time we had high load on our products since moving to 3 nodes. Therefore, while this problem existed all along, we didn&rsquo;t notice it until our systems were under stress.</p>
<h2 id="next-steps">Next Steps</h2>
<p>To solve these issues more quickly in the future, we should have better alerting and monitoring around application failures. Specifically:</p>
<ul>
<li>Alerting when we have a high number of <code>503 Service Unavailable</code> responses. These almost always indicate an infrastructure-level error.</li>
<li>Alerting when we have high node CPU usage. While pod CPU usage hasn&rsquo;t been great signal for us (pods occasionally spike), high node usage almost always indicates a problem.</li>
</ul>
<h2 id="reach-out">Reach Out</h2>
<p>If any of this looks like it&rsquo;s up your alley or you wanna learn more about our mission, be sure to email us at <a href="mailto:contact@pennlabs.org">contact@pennlabs.org</a> or <a href="https://pennlabs.org/apply">apply to be a part of Labs</a>! We&rsquo;ve got some fantastic teams working on interesting problems with a direct impact on campus.</p>
]]></content></item><item><title>The Chicken Thigh Manifesto</title><link>https://pawa.lt/posts/2021/07/the-chicken-thigh-manifesto/</link><pubDate>Tue, 06 Jul 2021 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2021/07/the-chicken-thigh-manifesto/</guid><description>I&amp;rsquo;ve tried to write this blog post many times before, but I always get stuck at the prose here. In the interest of actually writing this blog post, I&amp;rsquo;m going to skip it and get straight down to the good stuff. Maybe I&amp;rsquo;ll come back and write a story here. I probably will not.
Intro The key idea of these thighs is that we want to end up with bulk pre-cooked chicken that can enjoyed by itself or as an ingredient to other meals.</description><content type="html"><![CDATA[<p>I&rsquo;ve tried to write this blog post many times before, but I always get stuck at the prose here. In the interest of actually writing this blog post, I&rsquo;m going to skip it and get straight down to the good stuff. Maybe I&rsquo;ll come back and write a story here. I probably will not.</p>
<h2 id="intro">Intro</h2>
<p>The key idea of these thighs is that we want to end up with bulk pre-cooked chicken that can enjoyed by itself or as an ingredient to other meals. For this, we have a few desirable qualities:</p>
<ul>
<li>
<p><strong>Tenderness:</strong> If we want thighs that can be re-cooked, they must be relatively tender after their first cook.</p>
<p>We&rsquo;ll achieve this by using thighs (high fat content) and by using an instant pot (locks in moisture).</p>
</li>
<li>
<p><strong>Base flavor:</strong> We want the thighs to have a flavor that will make them appealing to eat alone but will not overpower dishes.</p>
<p>Our garlic soy sauce will provide a subtle but helpful flavor boost to cut some of the &ldquo;chicken taste&rdquo; out of the thighs.</p>
</li>
<li>
<p><strong>Manageability:</strong> I only want to cook chicken 1-2 times a week. I should be able to make 8+ chicken thighs at once.</p>
<p>The instant pot is a one-size-fits-all tool for doing exactly this.</p>
</li>
<li>
<p><strong>Healthy:</strong> If I&rsquo;m going to eat these almost every day, they better not kill me.</p>
<p>While chicken thighs aren&rsquo;t as lowfat as breasts, they&rsquo;re still pretty good for you, and we aren&rsquo;t adding anything too egregious to them. If you&rsquo;re watching calories or want to pump your protein consumption, you can switch to breasts at the expense of initial tenderness and recookability.</p>
</li>
</ul>
<h2 id="ingredients--recipe">Ingredients &amp; Recipe</h2>

<style>
.ingredient-section {
  border: 0.5ch solid;
  border-radius: 1ch;
  float: left;
  padding-left: 1ch;
  padding-right: 1ch;
  margin: 1em;
}
.ingredient-group-name {
  margin-top: 0.25em;
  margin-bottom: 0.25em;
}
ul {
  margin-top: 0em;
}
</style>

<div class="ingredient-section">
<h3 class="ingredient-group-name">Marinade</h3>

<ul>
<li>1 cup soy sauce</li>
<li>1 cup orange juice</li>
<li>2 cloves Garlic</li>
<li>1/4 cup olive oil</li>
<li>Misc. italian herbs (oregano, basil, etc)</li>
<li>Pepper</li>
</ul>

</div>

<style>
.ingredient-section {
  border: 0.5ch solid;
  border-radius: 1ch;
  float: left;
  padding-left: 1ch;
  padding-right: 1ch;
  margin: 1em;
}
.ingredient-group-name {
  margin-top: 0.25em;
  margin-bottom: 0.25em;
}
ul {
  margin-top: 0em;
}
</style>

<div class="ingredient-section">
<h3 class="ingredient-group-name">Thighs</h3>

<ul>
<li>Chicken Thighs</li>
<li>Salt</li>
<li>Pepper</li>
</ul>

</div>
<br style="clear:both"/>
<h3 id="prepping">Prepping</h3>
<ol>
<li>
<p>Mince garlic and add all marinade ingredients to mixing bowl. The measurements on this are very inexact, so taste and test different combinations yourself! Just taste and adjust until you like it.</p>
<p><strong>Steps below are optional but highly recommended</strong></p>
</li>
<li>
<p>Dry thighs with a paper towel and season all sides with salt and pepper</p>
</li>
<li>
<p>Heat large pan on high heat with a neutral oil like canola until the oil is almost smoking.</p>
<p>Cast-iron or carbon steel are best here for their heat retention, but nonstick is perfectly fine.</p>
</li>
<li>
<p>Cook thighs until they develop a hard sear.</p>
<p>The texture from this sear is what will give the thighs texture, so don&rsquo;t skimp on this if you want to eat these solo!</p>
</li>
<li>
<p>Remove thighs, taking care to not remove the sear.</p>
<p>If you&rsquo;re in a nonstick, the thighs will come off with no problem, but if you&rsquo;re in a cast iron or nonstick, use your heavy-duty tools to make sure the sear comes off with the chicken. The worst feeling is building a great sear and it getting stuck to the pan.</p>
</li>
</ol>
<p>Once your thighs are seared, they should look something like this (peep the marinade in the back):</p>
<figure>
    <img src="/img/recipes/chicken-thighs/seared.jpg"
         alt="Seared chicken thighs"/> <figcaption>
            <p>Seared chicken thighs</p>
        </figcaption>
</figure>

<h4 id="veggie-interlude">Veggie interlude</h4>
<p>If you&rsquo;re looking to cook up some veggies for the week, the <a href="https://food52.com/blog/12331-how-to-make-sauce-out-of-your-pan-s-brown-bits-a-k-a-fond">fond</a> left in your searing pan will make for some amazing sautéed vegetables. Just cut up some veggies, throw them in with bit of olive oil, and wait for delicious veggies that go with any meal. I&rsquo;ll usually do something like the following:</p>
<ol>
<li>Add olive oil with diced onions, sliced mushrooms, sliced celery</li>
<li>Once onions are translucent, add sliced carrots and diced bell pepper</li>
<li>Once carrots have softened up, add sliced garlic, red pepper flakes, and salt and pepper to taste</li>
</ol>
<p>Wait until everything&rsquo;s combined and store for later use! I love this combination standalone or along with proteins like these chicken thighs.</p>
<h3 id="finalizing">Finalizing</h3>
<p>Now that we&rsquo;ve got thighs and marinade, add both to the instant pot and cook on high pressure for 3 minutes. If you didn&rsquo;t sear the thighs, do 5 minutes instead.</p>
<p>I&rsquo;ll typically put a thigh on rice, spoon some of the marinade from the pot over top and garnish with green onions. Here&rsquo;s what it looks like! These are actually breasts, but the thighs will look similar.</p>
<figure>
    <img src="/img/recipes/chicken-thighs/breasts.jpeg"
         alt="Breasts with sauce over rice garnished with green onions"/> <figcaption>
            <p>Breasts with sauce over rice garnished with green onions</p>
        </figcaption>
</figure>

<p>If you cooked up vegetables earlier, this is a great time to reuse them. Add those veggies on top of the chicken + rice for a full meal.</p>
<h3 id="storing">Storing</h3>
<p>Separate the thighs and marinade, and store them however you desire in your fridge. The thighs will keep for a little over a week, and the marinade will keep for a few weeks.</p>
<p>To reheat this meal, just throw a thigh + rice + a spoonful of marinade in a bowl. Pop it in the microwave for 1.5 minutes, and you&rsquo;ve got a hot meal ready to go.</p>
<h2 id="other-meals">Other meals</h2>
<p>For a twist on an omelette, try an <a href="/recipes/oyakomelette">Oyakomelette</a>! It adds a little body to an omelette to turn it from a light breakfast into a hearty lunch or dinner.</p>
<p><img src="/img/recipes/chicken-thighs/omelette.jpeg" alt="Oyakomelette"></p>
<p>If you&rsquo;re looking for a grilled chicken salad, just slice the thighs, sear them off in a pan, and <a href="/recipes/grilled-chicken-salad/">follow my recipe</a> (or don&rsquo;t; it&rsquo;s just a salad).</p>
<p><img src="/img/recipes/chicken-thighs/salad.jpeg" alt="Salad"></p>
<p>Finally, if you&rsquo;re looking for a greasy late-night meal, make yourself a quesadilla. Fry up diced chicken in olive oil, combine with cheese and carmelized onions, pack the filling into a tortilla, sear in a pan, and you&rsquo;re good to go. I have made more quesadillas with this stuff than I&rsquo;d like to count, and they never disappoint.</p>
<p><img src="/img/recipes/chicken-thighs/quesadillas.jpeg" alt="Quesadilla"></p>
<p>Don&rsquo;t let these ideas limit you, though! You can use these thighs wherever you would use chicken. I&rsquo;ve made fried rice, tacos, sandwiches, and more with this stuff.</p>
<h2 id="conclusion">Conclusion</h2>
<p>These thighs have served me very well, and I&rsquo;m sure they&rsquo;ll serve you as well. Just make sure to stay creative with it! These being pre-cooked gives you lots of extra creative time to explore whatever dishes you desire.</p>
]]></content></item><item><title>Icarus: Taming the Kubernetes Beast</title><link>https://pawa.lt/posts/2020/05/icarus-taming-the-kubernetes-beast/</link><pubDate>Wed, 27 May 2020 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2020/05/icarus-taming-the-kubernetes-beast/</guid><description>NOTE: This post was cross-posted from the Penn Labs blog! Check it out over there as well and stay around to see some of the cool work the team does: https://pennlabs.org/blog/building-icarus
As of recent, we&amp;rsquo;ve moved our infrastructure away from Dokku and onto Kubernetes. Especially for an organization with high turnover (nature of being a club), you have to carefully consider the organizational impacts of moving to a technology as complex as Kubernetes.</description><content type="html"><![CDATA[<blockquote>
<p>NOTE: This post was cross-posted from the Penn Labs blog! Check it out over there as well and stay around to see some of the cool work the team does:
<a href="https://pennlabs.org/blog/building-icarus">https://pennlabs.org/blog/building-icarus</a></p>
</blockquote>
<p>As of recent, we&rsquo;ve moved our infrastructure away from <a href="http://dokku.viewdocs.io/dokku/">Dokku</a> and onto <a href="https://kubernetes.io/">Kubernetes</a>. Especially for an organization with high turnover (nature of being a club), you have to carefully consider the organizational impacts of moving to a technology as complex as Kubernetes. This post should serve as a bit of insight into how we dealt with that challenge and built technology to solve it.</p>
<h2 id="background">Background</h2>
<p>Typically when organizations decide to adopt Kubernetes, their organizational structure takes one of two patterns (or a combination of both):</p>
<h3 id="pattern-1-embedded-sres">Pattern 1: Embedded SREs</h3>
<p>In this pattern, each team has one or more Site Reliability Engineers embedded in each team. These SREs are responsible for managing the deployment infrastructure and general architecture for the product. Since this model gives teams Kubernetes experts right in the teams, these teams are responsible for managing their own Kubernetes manifests and general deployment strategy.</p>
<p>This system works great for teams with complicated application architectures whose applications might require special features like persistent volumes, cross-datacenter availability, complex scheduling requirements, etc. Think teams running Cassandra clusters, Solr clusters, or database services. In this model, SREs on each team can fit the Kubernetes configs to each team&rsquo;s needs. The clear drawback, however, is that your number of SREs required scales linearly with your number of teams.</p>
<p>Because of the (relative) simplicity of Penn Labs&rsquo; applications and only having 2-3 SREs to service 5 teams, we determined that this pattern was not an option.</p>
<h3 id="pattern-2-platform-team">Pattern 2: Platform Team</h3>
<p>In this pattern, there is a single team (or organization) that is responsible for building a high-level platform that other teams can use to deploy their applications. Building a platform involves significant up-front engineering work and a strong team to build the automation and interfaces required in order to create a frictionless deployment experience.</p>
<p>This approach typically works best for applications with simple architectures. If you can restrict your applications to following similar architectures, you can build a platform that takes advantage of those shared architecture, abstracting away the common configuration required by your applications.</p>
<p>In Labs, all of our web applications follow similar architectures: a single frontend (usually React) and a monolithic backend (usually Django). Since we have a common, simple architecture and we don&rsquo;t have the personell to man every team with an SRE, we decided to go with the platform approach. This approach fit nicely into our organizational structure since we already have a platform team that manages our <a href="https://github.com/pennlabs/platform">authentication system</a>.</p>
<h2 id="platform-abstractions">Platform Abstractions</h2>
<p>There are quite a few tools that fit together to make this system work (Hashicorp Vault, Grafana, etc.), but for the purposes of this post, I&rsquo;m going to focus solely on the tool that developers actually use: Icarus.</p>
<p>Kubernetes configs are hard and complicated. Kubernetes aims to be able to support any containerized application architecture possible, so it exposes every knob you could possibly want to turn. In simple setups, as you can see below, this leads to nothing other than repetition and confusion:</p>
<div class="highlight-wrapper">
    
        <div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-yaml" data-lang="yaml">---
<span style="color:#75715e"># Source: icarus/templates/services.yaml</span>
<span style="color:#66d9ef">apiVersion</span>: v1
<span style="color:#66d9ef">kind</span>: Service
<span style="color:#66d9ef">metadata</span>:
  <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
<span style="color:#66d9ef">spec</span>:
  <span style="color:#66d9ef">type</span>: ClusterIP
  <span style="color:#66d9ef">ports</span>:
    - <span style="color:#66d9ef">port</span>: <span style="color:#ae81ff">80</span>
      <span style="color:#66d9ef">targetPort</span>: <span style="color:#ae81ff">80</span>
  <span style="color:#66d9ef">selector</span>:
    <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
---
<span style="color:#75715e"># Source: icarus/templates/deployments.yaml</span>
<span style="color:#66d9ef">apiVersion</span>: apps/v1
<span style="color:#66d9ef">kind</span>: Deployment
<span style="color:#66d9ef">metadata</span>:
  <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
  <span style="color:#66d9ef">namespace</span>: default
  <span style="color:#66d9ef">labels</span>:
    <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
<span style="color:#66d9ef">spec</span>:
  <span style="color:#66d9ef">replicas</span>: <span style="color:#ae81ff">1</span>
  <span style="color:#66d9ef">selector</span>:
    <span style="color:#66d9ef">matchLabels</span>:
      <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
  <span style="color:#66d9ef">template</span>:
    <span style="color:#66d9ef">metadata</span>:
      <span style="color:#66d9ef">labels</span>:
        <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
    <span style="color:#66d9ef">spec</span>:
      <span style="color:#66d9ef">containers</span>:
        - <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;worker&#34;</span>
          <span style="color:#66d9ef">image</span>: <span style="color:#e6db74">&#34;pennlabs/website:latest&#34;</span>
          <span style="color:#66d9ef">imagePullPolicy</span>: IfNotPresent
          <span style="color:#66d9ef">ports</span>:
            - <span style="color:#66d9ef">containerPort</span>: <span style="color:#ae81ff">80</span>
          <span style="color:#66d9ef">envFrom</span>:
            - <span style="color:#66d9ef">secretRef</span>:
                <span style="color:#66d9ef">name</span>: test-secret
---
<span style="color:#75715e"># Source: icarus/templates/ingresses.yaml</span>
<span style="color:#66d9ef">apiVersion</span>: networking.k8s.io/v1beta1
<span style="color:#66d9ef">kind</span>: Ingress
<span style="color:#66d9ef">metadata</span>:
  <span style="color:#66d9ef">name</span>: testsite-serve
  <span style="color:#66d9ef">namespace</span>: default
<span style="color:#66d9ef">spec</span>:
  <span style="color:#66d9ef">rules</span>:
    - <span style="color:#66d9ef">host</span>: <span style="color:#e6db74">&#34;pennlabs.org&#34;</span>
      <span style="color:#66d9ef">http</span>:
        <span style="color:#66d9ef">paths</span>:
          - <span style="color:#66d9ef">path</span>: <span style="color:#e6db74">&#34;/&#34;</span>
            <span style="color:#66d9ef">backend</span>:
              <span style="color:#66d9ef">serviceName</span>: testsite-serve
              <span style="color:#66d9ef">servicePort</span>: <span style="color:#ae81ff">80</span>
  <span style="color:#66d9ef">tls</span>:
    - <span style="color:#66d9ef">hosts</span>:
        - <span style="color:#e6db74">&#34;pennlabs.org&#34;</span>
      <span style="color:#66d9ef">secretName</span>: pennlabs-org-tls
---
<span style="color:#75715e"># Source: icarus/templates/certificates.yaml</span>
<span style="color:#66d9ef">apiVersion</span>: cert-manager.io/v1alpha2
<span style="color:#66d9ef">kind</span>: Certificate
<span style="color:#66d9ef">metadata</span>:
  <span style="color:#66d9ef">name</span>: pennlabs-org
  <span style="color:#66d9ef">annotations</span>:
    <span style="color:#66d9ef">&#34;helm.sh/resource-policy&#34;: </span>keep
<span style="color:#66d9ef">spec</span>:
  <span style="color:#66d9ef">secretName</span>: pennlabs-org-tls
  <span style="color:#66d9ef">dnsNames</span>:
  - <span style="color:#e6db74">&#34;pennlabs.org&#34;</span>
  - <span style="color:#e6db74">&#34;*.pennlabs.org&#34;</span>
  <span style="color:#66d9ef">issuerRef</span>:
    <span style="color:#66d9ef">name</span>: wildcard-letsencrypt-prod
    <span style="color:#66d9ef">kind</span>: ClusterIssuer
    <span style="color:#66d9ef">group</span>: cert-manager.io</code></pre></div>
    
</div>

<p>As mentioned earlier, in our applications we can make the following assumptions to make our lives easier:</p>
<ul>
<li>There will be exactly one container per application</li>
<li>That application, if exposed to the world, will speak HTTP and want to be secured with HTTPS</li>
<li>That application will use secrets synced into Kubernetes by our Vault secret sync job</li>
</ul>
<p>With just these three assumptions, we can radically simplify our required configuration to the following for this example:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-yaml" data-lang="yaml"><span style="color:#66d9ef">deploy_version</span>: <span style="color:#ae81ff">0.1.15</span>

<span style="color:#66d9ef">applications</span>:
  - <span style="color:#66d9ef">name</span>: serve
    <span style="color:#66d9ef">image</span>: pennlabs/website
    <span style="color:#66d9ef">secret</span>: test-secret
    <span style="color:#66d9ef">ingress</span>:
      <span style="color:#66d9ef">hosts</span>:
        - <span style="color:#66d9ef">host</span>: pennlabs.org
          <span style="color:#66d9ef">paths</span>: [<span style="color:#e6db74">&#39;/&#39;</span>]
</code></pre></div><p>Let&rsquo;s dive into how this is possible.</p>
<h3 id="helm">Helm</h3>
<p>At the root of the transformation from Icarus file to full Kubernetes manifest is a tool called Helm. Helm describes itself as a &ldquo;package manager for Kubernetes&rdquo;, but for the purposes of this article, think of it as a templating tool that happens to be good at deploying to Kubernetes.</p>
<p>Helm operates off the <a href="https://golang.org/pkg/text/template/">Go template package</a>. This package, much like Python&rsquo;s Jinja or Ruby&rsquo;s ERB, lets you define a template file, specifying where to put information the user provides. Take the example of our <a href="https://github.com/pennlabs/icarus/blob/master/templates/services.yaml">service template</a>. This template defines a basic service with parameters:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-yaml" data-lang="yaml"><span style="color:#66d9ef">apiVersion</span>: v1
<span style="color:#66d9ef">kind</span>: Service
<span style="color:#66d9ef">metadata</span>:
  <span style="color:#66d9ef">name</span>: {{ $app_id | quote }}
<span style="color:#66d9ef">spec</span>:
  <span style="color:#66d9ef">type</span>: {{ .svc_type }}
  <span style="color:#66d9ef">ports</span>:
    - <span style="color:#66d9ef">port</span>: {{ .port }}
      <span style="color:#66d9ef">targetPort</span>: {{ .port }}
  <span style="color:#66d9ef">selector</span>:
    <span style="color:#66d9ef">name</span>: {{ $app_id | quote }}
</code></pre></div><p>Helm lets you define default values, so in Icarus, we set <code>.svc_type</code> to be <code>ClusterIP</code> by default and <code>.port</code> to be 80 by default. We also set <code>$app_id</code> to be <code>&lt;repository_name&gt;-&lt;application_name&gt;</code>. Applying these rules, we get the resulting configuration that we see above:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-yaml" data-lang="yaml"><span style="color:#75715e"># Source: icarus/templates/services.yaml</span>
<span style="color:#66d9ef">apiVersion</span>: v1
<span style="color:#66d9ef">kind</span>: Service
<span style="color:#66d9ef">metadata</span>:
  <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
<span style="color:#66d9ef">spec</span>:
  <span style="color:#66d9ef">type</span>: ClusterIP
  <span style="color:#66d9ef">ports</span>:
    - <span style="color:#66d9ef">port</span>: <span style="color:#ae81ff">80</span>
      <span style="color:#66d9ef">targetPort</span>: <span style="color:#ae81ff">80</span>
  <span style="color:#66d9ef">selector</span>:
    <span style="color:#66d9ef">name</span>: <span style="color:#e6db74">&#34;testsite-serve&#34;</span>
</code></pre></div><p>Icarus works simply by applying substitutions like this to create the final Kubernetes manifests, and then we use the DigitalOcean API to automate the deployment of these manifests to our actual cluster. I&rsquo;m glossing over a lot of detail here, but feel free to check out our <a href="https://github.com/pennlabs/orb-helm-tools/blob/master/src/commands/deploy.yml">deploy orb</a> for the specifics.</p>
<h2 id="end-result">End Result</h2>
<p>There&rsquo;s been a lot of talk here about organizational philosophy and YAML templating, but let&rsquo;s circle back to what this system actually gives us.</p>
<p>When one of our developers wants to create a new application, they can, completely independently, create a Git repo, add in a small CI config file, create their Icarus file, and add their proper secrets into Vault. Once that&rsquo;s done, they can Git push up, and their application will be published to the domain of their choice in just a few minutes. If you ask me, that&rsquo;s pretty cool.</p>
<h2 id="reach-out">Reach Out</h2>
<p>If any of this looks like it&rsquo;s up your alley or you wanna learn more about our mission, be sure to contact us at <a href="mailto:contact@pennlabs.org">contact@pennlabs.org</a> or <a href="https://pennlabs.org/apply">apply to be a part of Labs</a>! We&rsquo;ve got some fantastic teams working on interesting problems with a direct impact on campus.</p>
]]></content></item><item><title>Building a Kudos Button</title><link>https://pawa.lt/posts/2020/04/building-a-kudos-button/</link><pubDate>Sun, 12 Apr 2020 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2020/04/building-a-kudos-button/</guid><description>NOTE: FaunaDB the company died, so this approach no longer is in use. I work for Modal now, so I re-implemented the backend using Modal. Sadly I lost all my data, so please give me kudos if you like the articles!
When diving into WebRTC recently, I ran into this great article on the limitations of WebRTC, particularly related to its unreliability in doing NAT traversal. At the end of the article, I saw that it had this neat &amp;ldquo;kudos&amp;rdquo; button that when I hovered over it, upped the kudos count for the post.</description><content type="html"><![CDATA[<blockquote>
<p><strong>NOTE:</strong> FaunaDB the company died, so this approach no longer is in use. I work for <a href="https://modal.com">Modal</a> now, so I <a href="https://github.com/pawalt/personal-site/blob/806e2c70a086b39ba3609713b2159c2cd150fd40/modal_kudos.py">re-implemented the backend</a> using Modal. Sadly I lost all my data, so please give me kudos if you like the articles!</p>
</blockquote>
<p>When diving into WebRTC recently, I ran into <a href="http://blog.alexfreska.com/webrtc-not-quite-magic">this great article</a> on the limitations of WebRTC, particularly related to its unreliability in doing NAT traversal. At the end of the article, I saw that it had this neat &ldquo;kudos&rdquo; button that when I hovered over it, upped the kudos count for the post. It turns out that kudos are a feature of the <a href="https://svbtle.com/">Svtle</a> platform.</p>
<p>While I&rsquo;m not interested in moving my site over to Svtle, I wanted that button, so I decided to make it. The button is comprised of 3 main parts:</p>
<ol>
<li>CSS animation for the expanding/contracting and the color change</li>
<li>Netlify serverless function for tracking kudos count</li>
<li>Client-side JavaScript glue to hook up the animation to the serverless function</li>
</ol>
<h2 id="css">CSS</h2>
<p>When I started coding (2013ish), CSS quickly became my mortal enemy. Nothing made sense to me, and the rules felt very inconsistent. So much so, in fact, that I quickly gave up on writing my own CSS and just used Bootstrap (now Bulma) for every project moving forward. Naturally, this had me dreading working with CSS for this project.</p>
<figure>
    <img src="/img/howtocss.png"
         alt="An infrastructure guy tries to use CSS (2020, colorized)"/> <figcaption>
            <p>An infrastructure guy tries to use CSS (2020, colorized)</p>
        </figcaption>
</figure>

<p>Once I got past some initial bumps, though, I found that it turns out that 7 years of CSS improvement and Sass really have made CSS much friendlier. The way the final animation works is roughly:</p>
<ol>
<li>Start the dot off as a small size, <code>$dotSize</code></li>
<li>On hover, fill out to the size of the ring, <code>$filledSize</code>. Animate just using the <code>transition: 1s</code> attribute.</li>
<li>After 1s of hovering, run some JavaScript (timing method described below) to change the ring and circle classes to <code>filled</code>, which gives them a green color and the circle a slightly smaller size than the ring. Transition time <code>0.2s</code> felt right for this.</li>
</ol>
<p>These steps created a nice pop out and back in that I liked, so now I could move on to writing my server-side functions.</p>
<h2 id="netlify-functions">Netlify Functions</h2>
<p>Before this, I had been hosting my site on Github Pages. Pages worked great for me since I had no dynamic content, but this button threw a wrench in that. Moving my site over was largely uninteresting (to Netlify&rsquo;s credit), so I&rsquo;ll just focus on the functions here.</p>
<p>For storing like counts, I use FaunaDB, a serverless database with a generous free tier. Fauna is super easy to set up with Netlify. I&rsquo;ll note how to bootstrap at the end of this article. In my Fauna collection, I store documents that contain the title of the article with a list of which client IDs have given it kudos. Example:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-json" data-lang="json">{
  <span style="color:#f92672">&#34;title&#34;</span>: <span style="color:#e6db74">&#34;automating-k3s-deployment-on-proxmox&#34;</span>,
  <span style="color:#f92672">&#34;users&#34;</span>: [
    <span style="color:#e6db74">&#34;bb344507-a78f-4c6d-a71d-362056a8e6d8&#34;</span>,
    <span style="color:#e6db74">&#34;e894ef53-7a4a-4b6b-b081-e28adba82a04&#34;</span>,
    <span style="color:#e6db74">&#34;755408c9-f671-4ff0-b577-30c8075f3936&#34;</span>,
    <span style="color:#e6db74">&#34;768839bf-39b6-432e-96f2-38d5e1aa575a&#34;</span>,
    <span style="color:#e6db74">&#34;8ce3057f-22bb-47d7-9c29-c2dad1205fa0&#34;</span>
  ]
}
</code></pre></div><p>Once I hooked up Fauna, I simply created two serverless functions:</p>
<ul>
<li><code>get-kudos.js</code> - This returns a JSON response with the number of kudos and whether or not the current user has given the post kudos. Example:
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-json" data-lang="json">{
  <span style="color:#f92672">&#34;numClicked&#34;</span>: <span style="color:#ae81ff">5</span>,
  <span style="color:#f92672">&#34;userClicked&#34;</span>: <span style="color:#66d9ef">true</span>
}
</code></pre></div></li>
<li><code>add-kudos.js</code> - This adds the current user to the list of users who have given kudos to a post.</li>
</ul>
<h2 id="client-side-javascript">Client-side JavaScript</h2>
<p>To keep track of which users had liked which post, I generated UUIDs according to the UUIDv4 spec and stored them in local storage. At page load, I simply make an AJAX request to my <code>get-kudos</code> function to check if the logged in user has given the page kudos. If they have, I immediately add the <code>filled</code> CSS class to the ring and circle.</p>
<p>For keeping track of whether or not the user has hovered in the circle for 1 second, I used the <code>mouseenter</code> and <code>mouseleave</code> events. I have a variable <code>lastTime</code> which starts of as 0. When you enter the circle, <code>lastTime</code> gets set to the current UNIX time and a 1 second delayed function gets dispatched. When you leave the circle, <code>lastTime</code> gets set to 0.</p>
<p>After one second of entering the circle, the aforementioned delayed function is run which checks if <code>lastTime</code> is 1 second or more before the current time. If it is, the <code>add-kudos</code> request is sent off, and the ring and circle becomes filled!</p>
<p>This completes the code, creating the kudos effect:</p>
<figure class="center">
<img src="/img/kudos.gif" style="height: 10ch; width: auto">
</figure>
<h2 id="on-your-site">On Your Site</h2>
<p>This button is pretty self-contained, so it shouldn&rsquo;t be too hard to add to your own site if you&rsquo;re hosted on Netlify. Here are some rough guidelines:</p>
<ol>
<li>Copy over <a href="https://github.com/pawalt/personal-site/tree/master/functions">my functions directory</a></li>
<li>Hook up FaunaDB and rename your database to <code>personal-site</code>
<pre><code>$ netlify addons:create fauna
$ netlify addons:auth fauna
# now go in and change your database to personal-site
$ cd functions
$ npm install
$ netlify dev:exec node createdb.js
</code></pre></li>
<li>Copy in my <a href="https://github.com/pawalt/personal-site/blob/master/assets/js/kudos.js">kudos JavaScript</a> and <a href="https://github.com/pawalt/personal-site/blob/master/assets/scss/_kudos.scss">kudos SCSS</a> into your project and import them into your page</li>
<li>Add a <a href="https://github.com/pawalt/personal-site/blob/7b70ab729465b44429c7345d8d4ce9631826772a/layouts/posts/single.html#L44">kudos container</a> to your post template</li>
<li>Make sure your <code>netlify.toml</code> is set to <a href="https://github.com/pawalt/personal-site/blob/a83453a7393b37d649df4b4b3c1aced178e3733f/netlify.toml#L3"><code>npm install</code></a> for your functions.</li>
</ol>
<h2 id="conclusion">Conclusion</h2>
<p>After not doing any web development for so long, it was good to stretch those muscles a bit. With Hugo and Netlify, this process required basically no infrastructure work on my part, which was exactly what I was looking for.</p>
<p>Hopefully this was helpful, and if you enjoyed this article, feel free to give me some kudos :)</p>
]]></content></item><item><title>Automating k3s Deployment on Proxmox</title><link>https://pawa.lt/posts/2019/07/automating-k3s-deployment-on-proxmox/</link><pubDate>Fri, 26 Jul 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/07/automating-k3s-deployment-on-proxmox/</guid><description>For about 2 years now, I&amp;rsquo;ve been very interested in learning Kubernetes and potentially using it in my homelab, but there&amp;rsquo;s always been one main thing holding me back: the complexity. To get Kubernetes up the first time, I had to spin up a bunch of VMs manually and then put tons of binaries on each one of them. Even if I could get a cluster spun up, there was no chance I could reproduce my work or deploy anything useful on the cluster.</description><content type="html"><![CDATA[<p>For about 2 years now, I&rsquo;ve been very interested in learning Kubernetes and potentially using it in my homelab, but there&rsquo;s always been one main thing holding me back: the complexity. To get Kubernetes up the first time, I had to spin up a bunch of VMs manually and then put tons of binaries on each one of them. Even if I could get a cluster spun up, there was no chance I could reproduce my work or deploy anything useful on the cluster.</p>
<p>All of this changed when I found out about <a href="https://github.com/rancher/k3s">k3s</a>. k3s allows you to run a full Kubernetes-compliant cluster without &ldquo;a PhD in k8s clusterology&rdquo;. Its main innovation is packaging all the normal Kubernetes components into a single binary and refactoring some authentication to make it much easier for nodes to join the cluster.</p>
<p>Combined with some Ansible and Terraform knowledge, I decided to get cooking on some automation to make spinning up a new cluster a simple command.</p>
<h2 id="prerequisites">Prerequisites</h2>
<p>There are a few things I&rsquo;m not going to cover in this guide that you need to get started:</p>
<ul>
<li>Install Terraform
<ul>
<li>Just download Terraform from <a href="https://www.terraform.io/downloads.html">the downloads page</a> and drop it somewhere in your <code>$PATH</code></li>
</ul>
</li>
<li>Install the <a href="https://github.com/Telmate/terraform-provider-proxmox">Terraform Proxmox provider</a>
<ul>
<li>Use the <code>go install</code> command found in the README</li>
<li>Drop the binaries into <code>~/.terraform.d/plugins/</code></li>
</ul>
</li>
<li>Have a Proxmox host</li>
<li>Install <a href="https://docs.ansible.com/ansible/latest/installation_guide/intro_installation.html">Ansible</a></li>
</ul>
<h2 id="building-a-cloud-init-ubuntu-template">Building a cloud-init Ubuntu template</h2>
<p>In order for Terraform to work smoothly with new VMs, we need to make a template VM that has <a href="https://cloud-init.io/">cloud-init</a> on it. Cloud-init adds some packages to the VM that makes automatic provisioning possible. Thankfully, Proxmox has pretty good support for it.</p>
<p>SSH into your Proxmox host, and enter these commmands to create an Ubuntu VM with cloud-init:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ wget https://cloud-images.ubuntu.com/bionic/current/bionic-server-cloudimg-amd64.img
<span style="color:#75715e"># Export whatever storage pool you want to use to hold your VM images. Mine just happens to be named hermes_data</span>
$ export STORAGE_POOL<span style="color:#f92672">=</span>hermes_data
$ qm create <span style="color:#ae81ff">8000</span> --memory <span style="color:#ae81ff">2048</span> --net0 virtio,bridge<span style="color:#f92672">=</span>vmbr0
$ qm importdisk <span style="color:#ae81ff">8000</span> bionic-server-cloudimg-amd64.img $STORAGE_POOL
Formatting <span style="color:#e6db74">&#39;/data/images/8000/vm-8000-disk-0.raw&#39;</span>, fmt<span style="color:#f92672">=</span>raw size<span style="color:#f92672">=</span><span style="color:#ae81ff">2361393152</span>
    <span style="color:#f92672">(</span>100.00/100%<span style="color:#f92672">)</span>
$ qm set <span style="color:#ae81ff">8000</span> --scsihw virtio-scsi-pci --scsi0 $STORAGE_POOL:8000/vm-8000-disk-0.raw
update VM 8000: -scsi0 hermes_data:8000/vm-8000-disk-0.raw -scsihw virtio-scsi-pci
$ qm set <span style="color:#ae81ff">8000</span> --name ubuntu-ci
update VM 8000: -name ubuntu-ci
</code></pre></div><p>Now we have a VM with all the appropriate cloud-init packages installed. Now we just have to attach the hardware that cloud-init requires, and we can template the VM.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ qm set <span style="color:#ae81ff">8000</span> --ide2 $STORAGE_POOL:cloudinit
update VM 8000: -ide2 hermes_data:cloudinit
Formatting <span style="color:#e6db74">&#39;/data/images/8000/vm-8000-cloudinit.qcow2&#39;</span>, fmt<span style="color:#f92672">=</span>qcow2 size<span style="color:#f92672">=</span><span style="color:#ae81ff">4194304</span> cluster_size<span style="color:#f92672">=</span><span style="color:#ae81ff">65536</span> preallocation<span style="color:#f92672">=</span>metadata lazy_refcounts<span style="color:#f92672">=</span>off refcount_bits<span style="color:#f92672">=</span><span style="color:#ae81ff">16</span>
$ qm set <span style="color:#ae81ff">8000</span> --boot c --bootdisk scsi0
update VM 8000: -boot c -bootdisk scsi0
$ qm set <span style="color:#ae81ff">8000</span> --serial0 socket --vga serial0
update VM 8000: -serial0 socket -vga serial0
$ qm template <span style="color:#ae81ff">8000</span>
</code></pre></div><p>There we go! We now have a template VM that we can build our k3s nodes off of.</p>
<h2 id="deploying-the-vms">Deploying the VMs</h2>
<p>Now, on whatever machine you have Ansible and Terraform installed on, clone down my <code>proxmox-k3s</code> repo:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ git clone https://github.com/Pwpon500/proxmox-k3s
Cloning into <span style="color:#e6db74">&#39;proxmox-k3s&#39;</span>...
remote: Enumerating objects: 64, <span style="color:#66d9ef">done</span>.
remote: Counting objects: 100% <span style="color:#f92672">(</span>64/64<span style="color:#f92672">)</span>, <span style="color:#66d9ef">done</span>.
remote: Compressing objects: 100% <span style="color:#f92672">(</span>38/38<span style="color:#f92672">)</span>, <span style="color:#66d9ef">done</span>.
remote: Total <span style="color:#ae81ff">64</span> <span style="color:#f92672">(</span>delta 3<span style="color:#f92672">)</span>, reused <span style="color:#ae81ff">64</span> <span style="color:#f92672">(</span>delta 3<span style="color:#f92672">)</span>, pack-reused <span style="color:#ae81ff">0</span>
Unpacking objects: 100% <span style="color:#f92672">(</span>64/64<span style="color:#f92672">)</span>, <span style="color:#66d9ef">done</span>.
$ cd proxmox-k3s/proxmox-tf/prod
</code></pre></div><p>Now, edit the <code>main.tf</code> file to reflect the IPs of all your nodes as well as what SSH keys you want to use. Next, export the appropriate variables so that the Proxmox Terraform provider can connect to the Proxmox API:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ export PM_API_URL<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;https://&lt;node_ip&gt;:8006/api2/json&#34;</span>&lt;Paste&gt;
$ export PM_USER<span style="color:#f92672">=</span>root@pam
$ export PM_PASS<span style="color:#f92672">=</span>&lt;your_pass_here&gt;
</code></pre></div><p>If you don&rsquo;t set any of these, you&rsquo;ll be prompted for them whenever you do a <code>terraform plan</code> or a <code>terraform apply</code>.</p>
<p>Without any further ado, let&rsquo;s create some VMs!</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ terraform init
$ terraform plan
$ terraform apply
</code></pre></div><p>These commands will have a lot of output, but I don&rsquo;t want to paste it all here. It should be clear if things have worked.</p>
<p>Wait a few minutes for the VMs to finish doing their cloud-init inital configuration, and continue to the next step.</p>
<h2 id="applying-k3s-configs">Applying k3s Configs</h2>
<p>Now, go into the <code>ansible-roles</code> directory. All you need to edit here is the <code>inventory.toml</code> file. Change the IPs to whatever yours are, and add as many hosts as you created. The only important thing to remember is to have exactly 1 master node and to make the rest workers.</p>
<p>Once you&rsquo;ve entered in the appropriate IPs, apply your playbook:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ ansible-playbook -i inventory.toml playbook.yml -u ubuntu
</code></pre></div><p>This command will take a few minutes to complete, but if it finishes successfully, you have a fully working k3s cluster!</p>
<h2 id="testing-out-the-cluster">Testing Out the Cluster</h2>
<p>To test out the cluster, we&rsquo;re going to go into the master node and use the in-built <code>k3s kubectl</code>:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ ssh ubuntu@172.30.100.80
ubuntu@k3s-node-0:~$ sudo k3s kubectl get nodes
NAME         STATUS   ROLES    AGE   VERSION
k3s-node-0   Ready    master   2d    v1.14.4-k3s.1
k3s-node-1   Ready    worker   2d    v1.14.4-k3s.1
k3s-node-2   Ready    worker   2d    v1.14.4-k3s.1
k3s-node-3   Ready    worker   2d    v1.14.4-k3s.1
</code></pre></div><p>Looks pretty good. We have all the nodes properly joined up. Let&rsquo;s start a pod just to test things out:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ sudo k3s kubectl run lab-nginx --image<span style="color:#f92672">=</span>nginx --port<span style="color:#f92672">=</span><span style="color:#ae81ff">80</span>
kubectl run --generator<span style="color:#f92672">=</span>deployment/apps.v1 is DEPRECATED and will be removed in a future version. Use kubectl run --generator<span style="color:#f92672">=</span>run-pod/v1 or kubectl create instead.
deployment.apps/lab-nginx created
$ sudo k3s kubectl expose deployment lab-nginx --type<span style="color:#f92672">=</span>NodePort
service/lab-nginx exposed
$ sudo k3s kubectl port-forward svc/lab-nginx 8080:80 &amp;
$ curl localhost:8080
&lt;!DOCTYPE html&gt;
&lt;html&gt;
&lt;head&gt;
&lt;title&gt;Welcome to nginx!&lt;/title&gt;
&lt;style&gt;
    body <span style="color:#f92672">{</span>
        width: 35em;
        margin: <span style="color:#ae81ff">0</span> auto;
        font-family: Tahoma, Verdana, Arial, sans-serif;
    <span style="color:#f92672">}</span>
&lt;/style&gt;
&lt;/head&gt;
&lt;body&gt;
&lt;h1&gt;Welcome to nginx!&lt;/h1&gt;
&lt;p&gt;If you see this page, the nginx web server is successfully installed and
working. Further configuration is required.&lt;/p&gt;

&lt;p&gt;For online documentation and support please refer to
&lt;a href<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;http://nginx.org/&#34;</span>&gt;nginx.org&lt;/a&gt;.&lt;br/&gt;
Commercial support is available at
&lt;a href<span style="color:#f92672">=</span><span style="color:#e6db74">&#34;http://nginx.com/&#34;</span>&gt;nginx.com&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Thank you <span style="color:#66d9ef">for</span> using nginx.&lt;/em&gt;&lt;/p&gt;
&lt;/body&gt;
&lt;/html&gt;
</code></pre></div><p>Woohoo! Looks to me like this thing&rsquo;s working. If you want to get out the kubeconfig to run <code>kubectl</code> from your local machine, it&rsquo;s in <code>/etc/rancher/k3s/k3s.yaml</code>.</p>
<p>Now if you don&rsquo;t want it anymore, you can do a simple <code>terraform destroy</code> in the <code>proxmox-tf</code> directory, and you&rsquo;re back to where you started.</p>
<h2 id="conclusion">Conclusion</h2>
<p>Obviously, what we&rsquo;ve done here is exciting because we get a Kubernetes cluster with a pretty simple process. It&rsquo;s also crazy useful that we can now spin up entire VM clusters with just a simple command. Before, I had to do this through the Proxmox Web UI, but not anymore.</p>
<p>I hope all this was helpful! I had a ton of fun with this, and there&rsquo;s more of this kind of material coming in the future.</p>
]]></content></item><item><title>Caplance Development Update 3</title><link>https://pawa.lt/posts/2019/07/caplance-development-update-3/</link><pubDate>Sun, 07 Jul 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/07/caplance-development-update-3/</guid><description>🎉 We made it! 🎉 After 6 months of work, Caplance is finally at MVP. The brief release notes can be found here.
To get to MVP, I had to implement the following functionality:
Logging Config file parsing Combining all the previous updates, we get the following feature set for the MVP:
Packet listening Packet forwarding over UDP Direct reply from backends, allowing the load balancer to only have to handle incoming traffic Dynamic backend registration and deregistration Backend commands and health checks Logging Config file parsing I&amp;rsquo;ll briefly cover the improvements so far and then give a short demo of Caplance working.</description><content type="html"><![CDATA[<p>🎉 We made it! 🎉 After 6 months of work, Caplance is finally at MVP. The brief release notes can be found <a href="https://github.com/Pwpon500/caplance/releases/tag/v0.1.0">here</a>.</p>
<p>To get to MVP, I had to implement the following functionality:</p>
<ul>
<li>Logging</li>
<li>Config file parsing</li>
</ul>
<p>Combining all the previous updates, we get the following feature set for the MVP:</p>
<ul>
<li>Packet listening</li>
<li>Packet forwarding over UDP</li>
<li>Direct reply from backends, allowing the load balancer to only have to handle incoming traffic</li>
<li>Dynamic backend registration and deregistration</li>
<li>Backend commands and health checks</li>
<li>Logging</li>
<li>Config file parsing</li>
</ul>
<p>I&rsquo;ll briefly cover the improvements so far and then give a short demo of Caplance working.</p>
<h2 id="logging">Logging</h2>
<p>For logging, I decided to use the <a href="https://github.com/sirupsen/logrus">Logrus</a> package. It&rsquo;s frequently used by other packages, and it was super easy to switch to. I had already implemented logging with the <code>log</code> package from the go stdlib, so by importing <code>log github.com/sirupsen/logrus</code> in its place, I got instant compatibility. From there, I changed the log levels as appropriate. I&rsquo;m using the following log levels:</p>
<ul>
<li>Debug</li>
<li>Info</li>
<li>Error</li>
<li>Fatal</li>
<li>Panic</li>
</ul>
<p>I&rsquo;d like to eventually get away from panic and just call a graceful stop function with an error code, but that&rsquo;s nitpicking pretty hard considering where Caplance is right now. Here&rsquo;s a little demo of what the logging looks like:</p>
<p>Load balancer:</p>
<p><img src="/img/caplance_server_logging_demo.png" alt="caplance server logging demo"></p>
<p>Backend:</p>
<p><img src="/img/caplance_client_logging_demo.png" alt="caplance client logging demo"></p>
<h2 id="config-file-parsing">Config File Parsing</h2>
<p>For parsing configuration files, I used <a href="https://github.com/spf13/viper">Viper</a>. Viper allows me to use JSON, HCL, TOML, or YAML to write my config files so long as they follow my defined structure. The structure for configuration files follows this struct:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">type</span> <span style="color:#a6e22e">config</span> <span style="color:#66d9ef">struct</span> {
    <span style="color:#a6e22e">Client</span> <span style="color:#66d9ef">struct</span> {
        <span style="color:#a6e22e">ConnectIP</span> <span style="color:#66d9ef">string</span>
        <span style="color:#a6e22e">DataIP</span>    <span style="color:#66d9ef">string</span>
        <span style="color:#a6e22e">Name</span>      <span style="color:#66d9ef">string</span>
    }
    <span style="color:#a6e22e">Server</span> <span style="color:#66d9ef">struct</span> {
        <span style="color:#a6e22e">MngIP</span>           <span style="color:#66d9ef">string</span>
        <span style="color:#a6e22e">BackendCapacity</span> <span style="color:#66d9ef">int</span>
    }
    <span style="color:#a6e22e">VIP</span>  <span style="color:#66d9ef">string</span>
    <span style="color:#a6e22e">Test</span> <span style="color:#66d9ef">bool</span>

    <span style="color:#a6e22e">HealthRate</span>   <span style="color:#66d9ef">int</span>
    <span style="color:#a6e22e">ReadTimeout</span>  <span style="color:#66d9ef">int</span>
    <span style="color:#a6e22e">WriteTimeout</span> <span style="color:#66d9ef">int</span>

    <span style="color:#a6e22e">Sockaddr</span> <span style="color:#66d9ef">string</span>
}
</code></pre></div><p>The following defaults are also set:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetDefault</span>(<span style="color:#e6db74">&#34;Test&#34;</span>, <span style="color:#66d9ef">false</span>)
<span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetDefault</span>(<span style="color:#e6db74">&#34;HealthRate&#34;</span>, <span style="color:#ae81ff">20</span>)
<span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetDefault</span>(<span style="color:#e6db74">&#34;RegisterTimeout&#34;</span>, <span style="color:#ae81ff">10</span>)
<span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetDefault</span>(<span style="color:#e6db74">&#34;ReadTimeout&#34;</span>, <span style="color:#ae81ff">30</span>)
<span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetDefault</span>(<span style="color:#e6db74">&#34;WriteTimeout&#34;</span>, <span style="color:#ae81ff">10</span>)
<span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetDefault</span>(<span style="color:#e6db74">&#34;Sockaddr&#34;</span>, <span style="color:#e6db74">&#34;/var/run/caplance.sock&#34;</span>)
</code></pre></div><p>In YAML, an example config file looks like this:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-yaml" data-lang="yaml"><span style="color:#66d9ef">vip</span>: <span style="color:#ae81ff">10.0.0.50</span>

<span style="color:#66d9ef">client</span>:
  <span style="color:#66d9ef">dataIP</span>: <span style="color:#ae81ff">10.0.0.2</span>
  <span style="color:#66d9ef">name</span>: backend<span style="color:#ae81ff">-1</span>

<span style="color:#66d9ef">server</span>:
  <span style="color:#66d9ef">mngIP</span>: <span style="color:#ae81ff">10.0.0.1</span>
  <span style="color:#66d9ef">backendCapacity</span>: <span style="color:#ae81ff">53</span>
</code></pre></div><p>This config file accepts all given defaults, only filling in the required fields. I&rsquo;ll write more detailed documentation on the config file format soon, but for now, this will do. What&rsquo;s really cool about Viper is that I can just tell it where to go looking for config files, give it a struct to put data into, and it does everything else for me:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">readConfig</span>() {
    <span style="color:#66d9ef">if</span> <span style="color:#a6e22e">configLocation</span> <span style="color:#f92672">!=</span> <span style="color:#e6db74">&#34;&#34;</span> {
        <span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">SetConfigFile</span>(<span style="color:#a6e22e">configLocation</span>)
    }

    <span style="color:#a6e22e">err</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">ReadInConfig</span>()
    <span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">!=</span> <span style="color:#66d9ef">nil</span> {
        <span style="color:#a6e22e">log</span>.<span style="color:#a6e22e">Fatal</span>(<span style="color:#e6db74">&#34;Failed to read in config: &#34;</span> <span style="color:#f92672">+</span> <span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">Error</span>())
    }

    <span style="color:#a6e22e">conf</span> = <span style="color:#f92672">&amp;</span><span style="color:#a6e22e">config</span>{}
    <span style="color:#a6e22e">err</span> = <span style="color:#a6e22e">viper</span>.<span style="color:#a6e22e">Unmarshal</span>(<span style="color:#a6e22e">conf</span>)
    <span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">!=</span> <span style="color:#66d9ef">nil</span> {
        <span style="color:#a6e22e">log</span>.<span style="color:#a6e22e">Fatal</span>(<span style="color:#e6db74">&#34;Failed to unmarshal config into struct: &#34;</span> <span style="color:#f92672">+</span> <span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">Error</span>())
    }
}
</code></pre></div><p>Here, I&rsquo;m just telling Viper that if a specific config file location exists, use it. Then, I read in the config and unmarshal it into the <code>conf</code> struct for later use. Pretty cool stuff.</p>
<h1 id="demo">Demo</h1>
<p>Now for the fun part! I&rsquo;ve prepared a brief demo just to show that Caplance works as we expect it to.</p>
<p>In this demo, I&rsquo;ll have the following hosts:</p>
<table>
<thead>
<tr>
<th>Host</th>
<th>Description</th>
<th>IP</th>
</tr>
</thead>
<tbody>
<tr>
<td>h1</td>
<td>Load Balancer</td>
<td>10.0.0.1</td>
</tr>
<tr>
<td>h2</td>
<td>Backend 1</td>
<td>10.0.0.2</td>
</tr>
<tr>
<td>h3</td>
<td>Backend 2</td>
<td>10.0.0.3</td>
</tr>
<tr>
<td>h4</td>
<td>Backend 3</td>
<td>10.0.0.4</td>
</tr>
<tr>
<td>h5</td>
<td>Client</td>
<td>10.0.0.5</td>
</tr>
</tbody>
</table>
<p>The virtual IP for the cluster that the client will be calling to is 10.0.0.50. Let&rsquo;s start things up! I&rsquo;m only going to show output from h1 and h2 right now, but h2 and h3 are showing the same as h2.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h1 $ ../caplance server -f confs/simple_1.yaml
INFO<span style="color:#f92672">[</span>0000<span style="color:#f92672">]</span> Reading in config file
INFO<span style="color:#f92672">[</span>0000<span style="color:#f92672">]</span> Starting load balancer

h2 $ ../caplance client -f confs/simple_1.yaml
INFO<span style="color:#f92672">[</span>0000<span style="color:#f92672">]</span> Reading in config file
INFO<span style="color:#f92672">[</span>0000<span style="color:#f92672">]</span> Starting client
2019/07/07 12:42:39 rpc.Register: method <span style="color:#e6db74">&#34;Start&#34;</span> has <span style="color:#ae81ff">2</span> input parameters; needs exactly three
</code></pre></div><p>Now, I&rsquo;m going to start nginx on the backends and see what happens when I curl from h1!</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h2 $ nginx -p <span style="color:#e6db74">&#34;&#34;</span> -c confs/nginx_1.conf

h5 $ ./repeat_curl.sh <span style="color:#ae81ff">10</span>
Welcome to Onion Backend 2!
Welcome to Onion Backend 1!
Welcome to Onion Backend 2!
Welcome to Onion Backend 2!
Welcome to Onion Backend 1!
Welcome to Onion Backend 1!
Welcome to Onion Backend 2!
Welcome to Onion Backend 2!
Welcome to Onion Backend 3!
Welcome to Onion Backend 2!
</code></pre></div><p>Looks like it&rsquo;s working! In this limited example, it seems that we&rsquo;re probing 3 much less than the others, but that&rsquo;s just random chance. If we run this over many more iterations and count occurences out, we see things level out:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h5 $ ./repeat_curl.sh <span style="color:#ae81ff">1000</span> | sort | uniq -c
    <span style="color:#ae81ff">330</span> Welcome to Onion Backend 1!
    <span style="color:#ae81ff">336</span> Welcome to Onion Backend 2!
    <span style="color:#ae81ff">333</span> Welcome to Onion Backend 3!
</code></pre></div><p>Looks like things are working out! If you want to run this demo yourself, all the required files are <a href="https://github.com/Pwpon500/caplance/tree/master/demo">right in the repo</a>.</p>
<h1 id="conclusion">Conclusion</h1>
<p>This was by far the most difficult project I&rsquo;ve ever worked on, so having it done feels pretty surreal. Before signing off, I want to thank <a href="https://github.com/davish">Davis</a>, <a href="https://github.com/ArmaanT">Armaan</a>, and <a href="https://github.com/benleim">Ben</a> for listening to me rant and rant about Caplance. Having an ear to talk to is unimaginably helpful when working on a project like this.</p>
<p>Despite some of the language I&rsquo;ve been using, this is most certainly not the last work I&rsquo;ll be doing on Caplance. I still have much to learn and features I want to implement. Until then, however, thanks for reading, and I&rsquo;ll see you next time.</p>
]]></content></item><item><title>Caplance Development Update 2</title><link>https://pawa.lt/posts/2019/06/caplance-development-update-2/</link><pubDate>Sun, 30 Jun 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/06/caplance-development-update-2/</guid><description>I&amp;rsquo;m back! Over the past few weeks, I&amp;rsquo;ve had much more time to work on Caplance than I had in the prior months, so, naturally, a ton of work has gotten done in these few weeks. Specifically, I&amp;rsquo;ve implemented the following functionality:
Backend registration The following controls from the backends: PAUSE DEREGISTER RESUME HEALTH caplancectl to tell a running client process to issue one of those commands Graceful stop for backends Packet listening on NFQUEUE Refactor project structure If you&amp;rsquo;ve been keeping track, you&amp;rsquo;ll notice that this puts us very close to the Caplance MVP!</description><content type="html"><![CDATA[<p>I&rsquo;m back! Over the past few weeks, I&rsquo;ve had much more time to work on Caplance than I had in the prior months, so, naturally, a ton of work has gotten done in these few weeks. Specifically, I&rsquo;ve implemented the following functionality:</p>
<ul>
<li>Backend registration</li>
<li>The following controls from the backends:
<ul>
<li>PAUSE</li>
<li>DEREGISTER</li>
<li>RESUME</li>
<li>HEALTH</li>
</ul>
</li>
<li><code>caplancectl</code> to tell a running client process to issue one of those commands</li>
<li>Graceful stop for backends</li>
<li>Packet listening on NFQUEUE</li>
<li>Refactor project structure</li>
</ul>
<p>If you&rsquo;ve been keeping track, you&rsquo;ll notice that this puts us very close to the Caplance MVP! All that&rsquo;s left is config file parsing and some housekeeping.</p>
<p>Here are some of the more technically interesting parts of this revision of Caplance:</p>
<h2 id="nfqueue">NFQUEUE</h2>
<p>If you&rsquo;ve got some time and an appetite for some very interesting debugging, I highly recommend taking a look at <a href="https://pawa.lt/posts/2019/06/nfqueue-and-the-mysterious-reset/">my post about NFQUEUE</a>. If you don&rsquo;t, here&rsquo;s the TL;DR:</p>
<p>I switched how listening works in Caplance again. Now, I&rsquo;m using NFQUEUE, an iptables option to hand off packets from an iptables rule to a program in userspace before any packet processing is done.</p>
<h2 id="backend-communication">Backend Communication</h2>
<p>One of the things I wanted to learn from this project was how to write a system for keeping track of state between two hosts. I looked (extensively) into using RPC for this, but it just didn&rsquo;t seem like the right move. RPC is designed to make &ldquo;remote process calls&rdquo;, but I don&rsquo;t really need that. I need a way to keep track of state. I could&rsquo;ve mangled RPC to do this, but I think just writing it myself turned out to be more elegant.</p>
<p>The first step in creating this was creating the interface. I created a simple interface called <code>Communicator</code> to write data, read data, and close the underlying connection:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#75715e">// Communicator is the connection manager for a backend
</span><span style="color:#75715e"></span><span style="color:#66d9ef">type</span> <span style="color:#a6e22e">Communicator</span> <span style="color:#66d9ef">interface</span> {
    <span style="color:#a6e22e">ReadLine</span>() (<span style="color:#66d9ef">string</span>, <span style="color:#66d9ef">error</span>)
    <span style="color:#a6e22e">WriteLine</span>(<span style="color:#a6e22e">data</span> <span style="color:#66d9ef">string</span>) <span style="color:#66d9ef">error</span>
    <span style="color:#a6e22e">Close</span>() <span style="color:#66d9ef">error</span>
}
</code></pre></div><p>To actually implement this, I created a <code>TCPCommunicator</code> struct and filled in the methods.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#75715e">// TCPCommunicator is an implementation of BackendCommunicator over TCP
</span><span style="color:#75715e"></span><span style="color:#66d9ef">type</span> <span style="color:#a6e22e">TCPCommunicator</span> <span style="color:#66d9ef">struct</span> {
    <span style="color:#a6e22e">reader</span>       <span style="color:#f92672">*</span><span style="color:#a6e22e">bufio</span>.<span style="color:#a6e22e">Reader</span>
    <span style="color:#a6e22e">writer</span>       <span style="color:#f92672">*</span><span style="color:#a6e22e">bufio</span>.<span style="color:#a6e22e">Writer</span>
    <span style="color:#a6e22e">conn</span>         <span style="color:#a6e22e">net</span>.<span style="color:#a6e22e">Conn</span>
    <span style="color:#a6e22e">readTimeout</span>  <span style="color:#a6e22e">time</span>.<span style="color:#a6e22e">Duration</span>
    <span style="color:#a6e22e">writeTimeout</span> <span style="color:#a6e22e">time</span>.<span style="color:#a6e22e">Duration</span>
}
</code></pre></div><p>This struct has a reader and writer to get data from the underling connection. It also has a readTimeout and writeTimeout. While the writeTimeout is fairly mundane, the readTimeout turns out to be surprisingly useful.</p>
<p>In order to deregister a backend after a certain period of inactivity, I can just use the <code>SetReadDeadline(t time.Time) error</code> function with the <code>readTimeout</code>. Then, if the read errors out, and the error is a timeout, I know that the inactivity window has closed. Baking this logic into the TCPCommunicator, I now have a powerful tool that I can use on both the backend and load balancer side!</p>
<p>I also implemented the following commands that the client can send to the server:</p>
<ul>
<li>PAUSE - client requesting to stay registered but not have packets forwarded to it</li>
<li>DEREGISTER - client requesting to deregister</li>
<li>RESUME - client requesting to have packets forwarded after a PAUSE</li>
<li>HEALTH - message sent every few seconds to stop server from deregistering the client due to inactivity</li>
</ul>
<p>The way these are implemented isn&rsquo;t particularly interesting. I&rsquo;m just using a bufio reader to read data off the connection and parsing it with simple string tools.</p>
<h2 id="caplancectl">Caplancectl</h2>
<p>This isn&rsquo;t the most complex part of Caplance by any means, but I was blown away by how simple the <code>net/rpc</code> package made writing caplancectl, so I wanted to share.</p>
<p>Caplancectl connects to Caplance using a unix socket that Caplance creates at <code>/var/run/caplance.sock</code>. Caplance listens just as it would with a TCP listener:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">c</span>.<span style="color:#a6e22e">unixSock</span>, <span style="color:#a6e22e">err</span> = <span style="color:#a6e22e">net</span>.<span style="color:#a6e22e">Listen</span>(<span style="color:#e6db74">&#34;unix&#34;</span>, <span style="color:#a6e22e">SOCKADDR</span>)
</code></pre></div><p>Now that I&rsquo;ve got that socket, I can hand it off to the <code>rpc</code> package, and the rest of the networking is done for me!</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">rpc</span>.<span style="color:#a6e22e">Register</span>(<span style="color:#a6e22e">c</span>) <span style="color:#75715e">// registering the client to receive requests over rpc
</span><span style="color:#75715e"></span><span style="color:#a6e22e">rpc</span>.<span style="color:#a6e22e">HandleHTTP</span>()
<span style="color:#a6e22e">http</span>.<span style="color:#a6e22e">Serve</span>(<span style="color:#a6e22e">c</span>.<span style="color:#a6e22e">unixSock</span>, <span style="color:#66d9ef">nil</span>)
</code></pre></div><p>Now that I&rsquo;ve got the RPC listening, all I have to do is define some methods that the RPC can execute. The RPC registers any methods from the registered object that take two arguments for which the latter is a pointer and returns an error. More formally, the methods must have this signature:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> (<span style="color:#a6e22e">t</span> <span style="color:#f92672">*</span><span style="color:#a6e22e">T</span>) <span style="color:#a6e22e">MethodName</span>(<span style="color:#a6e22e">argType</span> <span style="color:#a6e22e">T1</span>, <span style="color:#a6e22e">replyType</span> <span style="color:#f92672">*</span><span style="color:#a6e22e">T2</span>) <span style="color:#66d9ef">error</span>
</code></pre></div><p>For example, my pause function looks like this:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#75715e">// Pause command from caplancectl
</span><span style="color:#75715e"></span><span style="color:#66d9ef">func</span> (<span style="color:#a6e22e">c</span> <span style="color:#f92672">*</span><span style="color:#a6e22e">Client</span>) <span style="color:#a6e22e">Pause</span>(<span style="color:#a6e22e">req</span> <span style="color:#f92672">*</span><span style="color:#66d9ef">string</span>, <span style="color:#a6e22e">reply</span> <span style="color:#f92672">*</span><span style="color:#66d9ef">string</span>) <span style="color:#66d9ef">error</span> {
    <span style="color:#a6e22e">err</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">c</span>.<span style="color:#a6e22e">pause</span>()
    <span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">==</span> <span style="color:#66d9ef">nil</span> {
        <span style="color:#f92672">*</span><span style="color:#a6e22e">reply</span> = <span style="color:#e6db74">&#34;Pause request sent&#34;</span>
    } <span style="color:#66d9ef">else</span> {
        <span style="color:#f92672">*</span><span style="color:#a6e22e">reply</span> = <span style="color:#e6db74">&#34;Pause request encountered an error: &#34;</span> <span style="color:#f92672">+</span> <span style="color:#a6e22e">err</span>.<span style="color:#a6e22e">Error</span>()
    }
    <span style="color:#66d9ef">return</span> <span style="color:#66d9ef">nil</span>
}
</code></pre></div><p>That&rsquo;s all that needs to be done from the server side! If you think that&rsquo;s easy, the client side is even easier.</p>
<p>On the client side, I can just use the RPC package to dial and then call whatever method I want. For example, if I wanted to issue a pause request, I would only need this code:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">client</span>, <span style="color:#a6e22e">err</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">rpc</span>.<span style="color:#a6e22e">DialHTTP</span>(<span style="color:#e6db74">&#34;unix&#34;</span>, <span style="color:#a6e22e">SOCKADDR</span>)
<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">!=</span> <span style="color:#66d9ef">nil</span> {
    <span style="color:#a6e22e">log</span>.<span style="color:#a6e22e">Fatal</span>(<span style="color:#a6e22e">err</span>)
}

<span style="color:#66d9ef">var</span> <span style="color:#a6e22e">reply</span> <span style="color:#66d9ef">string</span>
<span style="color:#a6e22e">err</span> = <span style="color:#a6e22e">client</span>.<span style="color:#a6e22e">Call</span>(<span style="color:#e6db74">&#34;Client.Pause&#34;</span>, <span style="color:#e6db74">&#34;&#34;</span>, <span style="color:#f92672">&amp;</span><span style="color:#a6e22e">reply</span>)
<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">!=</span> <span style="color:#66d9ef">nil</span> {
    <span style="color:#a6e22e">log</span>.<span style="color:#a6e22e">Fatal</span>(<span style="color:#a6e22e">err</span>)
}
<span style="color:#a6e22e">fmt</span>.<span style="color:#a6e22e">Println</span>(<span style="color:#a6e22e">reply</span>)
</code></pre></div><p>Pretty cool, right?</p>
<h1 id="future-work">Future Work</h1>
<p>I&rsquo;m pretty close to MVP. To get there, I have to implement the following:</p>
<ul>
<li>Config file parsing</li>
<li>Non-stdout logging</li>
</ul>
<p>I&rsquo;ve got some more features I&rsquo;d like to implement, but we&rsquo;re gonna focus on MVP for now. These shouldn&rsquo;t be too hard to get implemented, so keep you eyes peeled for update 3!</p>
]]></content></item><item><title>NFQUEUE and the Mysterious RESET</title><link>https://pawa.lt/posts/2019/06/nfqueue-and-the-mysterious-reset/</link><pubDate>Thu, 27 Jun 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/06/nfqueue-and-the-mysterious-reset/</guid><description>While working on my load balancer, Caplance, I ran into a very strange error when trying to establish a connection between a client and a backend. Before I go to deep, though, let me give a quick intro into how TCP connection establishment works.
Types of Message When establishing a TCP connection, there are 4 possible packet types you could see:
SYN (S in tcpdump) - Always the first message sent.</description><content type="html"><![CDATA[<p>While working on my load balancer, <a href="https://github.com/pwpon500/Caplance">Caplance</a>, I ran into a very strange error when trying to establish a connection between a client and a backend. Before I go to deep, though, let me give a quick intro into how TCP connection establishment works.</p>
<h3 id="types-of-message">Types of Message</h3>
<p>When establishing a TCP connection, there are 4 possible packet types you could see:</p>
<ul>
<li>SYN (<code>S</code> in <code>tcpdump</code>) - Always the first message sent. It is the message from the client to the server effectively saying &ldquo;I would like to connect.&rdquo;</li>
<li>SYN-ACK (<code>S.</code> in <code>tcpdump</code>) - The server&rsquo;s reply to a SYN if it wishes to accept the connection.</li>
<li>ACK (<code>A</code> or <code>.</code> in <code>tcpdump</code>) - The client&rsquo;s reply to a SYN-ACK, fully establishing the TCP connection.</li>
<li>RESET (<code>R</code> in <code>tcpdump</code>) - The server&rsquo;s response to a SYN if it wishes to refuse the connection.</li>
</ul>
<p>In a successful connection, the order will be SYN, SYN-ACK, ACK. In an unsuccessful connection, the order will be SYN, RESET.</p>
<h3 id="operating-systems-and-tcp">Operating Systems and TCP</h3>
<p>When a program wants to use TCP, it asks the OS to &ldquo;bind&rdquo; to a TCP port. Then, when packets come into that TCP port, the OS hands off the packets to the program. The program can then reply to those packets however it wants (accepting or rejecting the incoming connections).</p>
<p>If a program is not bound to a port and a SYN is sent to that port, the OS immediately replies to the SYN with a RESET.</p>
<p>In the case of Caplance, the load balancer isn&rsquo;t listening on a specific TCP port. Instead, it listens for any and all TCP packets. As a result, no bindings are created by Caplance.</p>
<h2 id="the-bug">The Bug</h2>
<p>Ok, now on to the fun part. I had finally gotten to the point in Caplance where I could send a request to the load balancer&rsquo;s IP and have a backend respond to the request, or so I thought. For the purposes of this demo, h1 is the load balancer, h2 is the backend, and h3 is the client. 10.0.0.50 is the virtual IP that the load balancer is serving. Let&rsquo;s start up a netcat TCP server on h2 and try to send a request from h3:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h2 $ nc -l 10.0.0.50 <span style="color:#ae81ff">8080</span>

h3 $ telnet 10.0.0.50 <span style="color:#ae81ff">8080</span>
Trying 10.0.0.50...
telnet: Unable to connect to remote host: Connection refused
</code></pre></div><p>Well, what can you expect really? Things never work on the first try. This time, let&rsquo;s do the exact same thing but use <code>tcpdump</code> to see what&rsquo;s happening on h2.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h2 $ tcpdump -i any host 10.0.0.50 and tcp -n
tcpdump: verbose output suppressed, use -v or -vv <span style="color:#66d9ef">for</span> full protocol decode
listening on any, link-type LINUX_SLL <span style="color:#f92672">(</span>Linux cooked<span style="color:#f92672">)</span>, capture size <span style="color:#ae81ff">262144</span> bytes
20:57:32.941669 IP 10.0.0.3.36040 &gt; 10.0.0.50.8080: Flags <span style="color:#f92672">[</span>S<span style="color:#f92672">]</span>, seq 252176171, win 29200, options <span style="color:#f92672">[</span>mss 1460,sackOK,TS val <span style="color:#ae81ff">2262321703</span> ecr 0,nop,wscale 9<span style="color:#f92672">]</span>, length <span style="color:#ae81ff">0</span>
20:57:32.941701 IP 10.0.0.50.8080 &gt; 10.0.0.3.36040: Flags <span style="color:#f92672">[</span>S.<span style="color:#f92672">]</span>, seq 492629165, ack 252176172, win 28960, options <span style="color:#f92672">[</span>mss 1460,sackOK,TS val <span style="color:#ae81ff">4192526147</span> ecr 2262321703,nop,wscale 9<span style="color:#f92672">]</span>, length <span style="color:#ae81ff">0</span>
</code></pre></div><p>Here, <code>tcpdump</code> is showing us that the backend is seeing 2 packets. One is a SYN from the client to the server, and one is a SYN-ACK from the server to the client. If you remember, these are exactly the messages we expect &hellip; minus an ACK from the client.</p>
<p>At this point, the logical conclusion is that packets are able to flow from client to server, but the ones going server to client aren&rsquo;t making it. Just for sanity, let&rsquo;s see what the <code>tcpdump</code> output looks like for the client.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h3 $ tcpdump -i any host 10.0.0.3 and tcp -n
tcpdump: verbose output suppressed, use -v or -vv <span style="color:#66d9ef">for</span> full protocol decode
listening on any, link-type LINUX_SLL <span style="color:#f92672">(</span>Linux cooked<span style="color:#f92672">)</span>, capture size <span style="color:#ae81ff">262144</span> bytes
21:11:45.055143 IP 10.0.0.3.37388 &gt; 10.0.0.50.8080: Flags <span style="color:#f92672">[</span>S<span style="color:#f92672">]</span>, seq 3009015791, win 29200, options <span style="color:#f92672">[</span>mss 1460,sackOK,TS val <span style="color:#ae81ff">2263173829</span> ecr 0,nop,wscale 9<span style="color:#f92672">]</span>, length <span style="color:#ae81ff">0</span>
21:11:45.059240 IP 10.0.0.50.8080 &gt; 10.0.0.3.37388: Flags <span style="color:#f92672">[</span>R.<span style="color:#f92672">]</span>, seq 0, ack 3009015792, win 0, length <span style="color:#ae81ff">0</span>
21:11:45.059248 IP 10.0.0.50.8080 &gt; 10.0.0.3.37388: Flags <span style="color:#f92672">[</span>S.<span style="color:#f92672">]</span>, seq 2629426318, ack 3009015792, win 28960, options <span style="color:#f92672">[</span>mss 1460,sackOK,TS val <span style="color:#ae81ff">4193378273</span> ecr 2263173829,nop,wscale 9<span style="color:#f92672">]</span>, length <span style="color:#ae81ff">0</span>
</code></pre></div><p>&hellip; What?</p>
<p>At this point, I&rsquo;m baffled. Let&rsquo;s recap what&rsquo;s going on here. The server is seeing a SYN and SYN-ACK. However, the client is seeing a SYN, SYN-ACK, and RESET. How is it possible that the client could be sent both a SYN-ACK and a RESET from the same source??</p>
<p>If you&rsquo;re interested in this stuff, I invite you to sit for a minute and think about why this might be happening. This is one of the more interesting bugs I&rsquo;ve run into in a while. The presence of a SYN-ACK and RESET concurrently really interested me.</p>
<h3 id="the-bug-revealed">The Bug Revealed</h3>
<p>Remember when I said that Caplance isn&rsquo;t binding to a specific port? That turns out to be the key to the puzzle here. <strong>Since Caplance listens for all TCP connections instead of a single port, it never sends a &ldquo;bind&rdquo; request to the OS. This means that the OS will still respond to all incoming TCP connections with a RESET.</strong></p>
<p>To clarify, let&rsquo;s think about the connection establishment journey:</p>
<ol>
<li>SYN leaves the client</li>
<li>SYN hits the load balancer
<ol>
<li>Because the load balancer isn&rsquo;t actually bound to a TCP port, the OS replies to the client with a RESET</li>
<li>The load balancer forwards the SYN to the backend</li>
</ol>
</li>
<li>SYN hits the backend
<ol>
<li>Backend replies to the client with SYN-ACK</li>
</ol>
</li>
<li>Client receives RESET</li>
<li>Client receives SYN-ACK</li>
</ol>
<p>As is typical, the actual bug is pretty simple in hindsight, but it can be hard to find in the moment.</p>
<h2 id="enter-nfqueue">Enter NFQUEUE</h2>
<p>So now we&rsquo;re in a conundrum: we want to stop the OS from seeing packets destined for the VIP, but we also want the OS to give packets destined to the VIP to our program so it can forward then to the appropriate backend. This is where NFQUEUE comes into play.</p>
<p>NFQUEUE is a iptables filter rule that queues up incoming packets onto a queue, waiting to be processed by some program in userspace. Let&rsquo;s see an example rule.</p>
<p>The following rules will drop all tcp and udp packets destined for the IP <code>10.0.0.50</code>:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ iptables -A INPUT -d 10.0.0.50/32 -p tcp -j DROP
$ iptables -A INPUT -d 10.0.0.50/32 -p udp -j DROP
</code></pre></div><p>What if instead, we want to put all tcp and udp destined for <code>10.0.0.50</code> onto a NFQUEUE with the id 0? We can simply modify our previous rule as follows:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">$ iptables -A INPUT -d 10.0.0.50/32 -p tcp -j NFQUEUE --queue-num <span style="color:#ae81ff">0</span>
$ iptables -A INPUT -d 10.0.0.50/32 -p udp -j NFQUEUE --queue-num <span style="color:#ae81ff">0</span>
</code></pre></div><p>Now, we have a single source (queue 0) that we can consume all our packets off of! There are bindings to do this in many languages, but I&rsquo;m using Go for Caplance, so here&rsquo;s a snippet of how this looks in Go.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">b</span>.<span style="color:#a6e22e">nfq</span>, <span style="color:#a6e22e">err</span> = <span style="color:#a6e22e">netfilter</span>.<span style="color:#a6e22e">NewNFQueue</span>(<span style="color:#ae81ff">0</span>, <span style="color:#ae81ff">100</span>, <span style="color:#a6e22e">netfilter</span>.<span style="color:#a6e22e">NF_DEFAULT_PACKET_SIZE</span>)
<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">!=</span> <span style="color:#66d9ef">nil</span> {
    <span style="color:#a6e22e">log</span>.<span style="color:#a6e22e">Panicln</span>(<span style="color:#a6e22e">err</span>)
}
<span style="color:#a6e22e">packetChan</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">b</span>.<span style="color:#a6e22e">nfq</span>.<span style="color:#a6e22e">GetPackets</span>()
<span style="color:#a6e22e">stopped</span> <span style="color:#f92672">:=</span> <span style="color:#66d9ef">false</span>
<span style="color:#66d9ef">for</span> !<span style="color:#a6e22e">stopped</span> {
    <span style="color:#66d9ef">select</span> {
    <span style="color:#66d9ef">case</span> <span style="color:#a6e22e">packet</span> <span style="color:#f92672">:=</span> <span style="color:#f92672">&lt;-</span><span style="color:#a6e22e">packetChan</span>:
        <span style="color:#a6e22e">b</span>.<span style="color:#a6e22e">packets</span> <span style="color:#f92672">&lt;-</span> <span style="color:#a6e22e">packet</span>.<span style="color:#a6e22e">Packet</span>.<span style="color:#a6e22e">Data</span>()
        <span style="color:#a6e22e">packet</span>.<span style="color:#a6e22e">SetVerdict</span>(<span style="color:#a6e22e">netfilter</span>.<span style="color:#a6e22e">NF_DROP</span>)
    <span style="color:#66d9ef">case</span> <span style="color:#a6e22e">sig</span> <span style="color:#f92672">:=</span> <span style="color:#f92672">&lt;-</span><span style="color:#a6e22e">b</span>.<span style="color:#a6e22e">stopChan</span>:
        <span style="color:#a6e22e">b</span>.<span style="color:#a6e22e">stopChan</span> <span style="color:#f92672">&lt;-</span> <span style="color:#a6e22e">sig</span>
        <span style="color:#a6e22e">stopped</span> = <span style="color:#66d9ef">true</span>
    }
}
</code></pre></div><p>Don&rsquo;t worry if you don&rsquo;t understand the Go specifics of this code. The important thing is that we&rsquo;re creating a NFQUEUE receiver called <code>b.nfq</code>. Then, we create a channel called <code>packetChan</code> off of which we can consume whatever packets come in to queue 0.</p>
<h3 id="running-it">Running It</h3>
<p>Let&rsquo;s make sure this works! If it does, we should see whatever we type into h3 popping up in the TCP server on h2.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash">h3 $ telnet 10.0.0.50 <span style="color:#ae81ff">8080</span>
Trying 10.0.0.50...
Connected to 10.0.0.50.
Escape character is <span style="color:#e6db74">&#39;^]&#39;</span>.
Is this working?
It is!!!!!
 ^<span style="color:#f92672">]</span>
telnet&gt; close
Connection closed.

h2 $ nc -l 10.0.0.50 <span style="color:#ae81ff">8080</span>
Is this working?
It is!!!!
</code></pre></div><p>We did it! NFQUEUE was the answer to our problems.</p>
<h2 id="conclusion">Conclusion</h2>
<p>When I found out about NFQUEUE, I was dumbstruck. It solved so many of my problems - creating a single source from which I could consume packets, stopping the load balancer from replying to the wrong packets, and overall just cleaning up my code. If you&rsquo;ve got a project like this, I highly recommend looking at NFQUEUE as a potential option.</p>
]]></content></item><item><title>My Neovim Go Setup</title><link>https://pawa.lt/posts/2019/06/my-neovim-go-setup/</link><pubDate>Sat, 01 Jun 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/06/my-neovim-go-setup/</guid><description>Recently, I moved from writing my Go in VSCode to writing it in Neovim. I did this for two main reasons:
Ubiquity
I use Neovim for writing pretty much all my other code (except Java), so it was weird to use VSCode for this one purpose.
Terminal Integration
I hate the VS Code terminal. It doesn&amp;rsquo;t have the character support I want, and it overrides a bunch of keys (for example, Ctrl+f for fish autocomplete).</description><content type="html"><![CDATA[<p>Recently, I moved from writing my Go in VSCode to writing it in Neovim. I did this for two main reasons:</p>
<ol>
<li>
<p>Ubiquity</p>
<p>I use Neovim for writing pretty much all my other code (except Java), so it was weird to use VSCode for this one purpose.</p>
</li>
<li>
<p>Terminal Integration</p>
<p>I <em>hate</em> the VS Code terminal. It doesn&rsquo;t have the character support I want, and it overrides a bunch of keys (for example, Ctrl+f for fish autocomplete). I&rsquo;m also in the terminal already when I&rsquo;m doing Go development, so being able to just type <code>vim</code> and instantly be in my editor is a blessing.</p>
</li>
</ol>
<p>Before I start, it&rsquo;s important to note that I&rsquo;m currently doing all this on Ubuntu 18.04 with Go 1.12, provided via the <code>longsleep/golang-backports</code> ppa. For more info on how to install Go, <a href="https://github.com/golang/go/wiki/Ubuntu">visit the wiki</a></p>
<h2 id="basics">Basics</h2>
<p>The first bit of configuration is my Go tab settings. These are standard for Go, and even if you don&rsquo;t set them, <code>gofmt</code> will enforce them anyway.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-vim" data-lang="vim"><span style="color:#a6e22e">au</span> <span style="color:#a6e22e">FileType</span> <span style="color:#a6e22e">go</span> <span style="color:#a6e22e">set</span> <span style="color:#a6e22e">noexpandtab</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">au</span> <span style="color:#a6e22e">FileType</span> <span style="color:#a6e22e">go</span> <span style="color:#a6e22e">set</span> <span style="color:#a6e22e">shiftwidth</span>=<span style="color:#ae81ff">4</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">au</span> <span style="color:#a6e22e">FileType</span> <span style="color:#a6e22e">go</span> <span style="color:#a6e22e">set</span> <span style="color:#a6e22e">softtabstop</span>=<span style="color:#ae81ff">4</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">au</span> <span style="color:#a6e22e">FileType</span> <span style="color:#a6e22e">go</span> <span style="color:#a6e22e">set</span> <span style="color:#a6e22e">tabstop</span>=<span style="color:#ae81ff">4</span><span style="color:#960050;background-color:#1e0010">
</span></code></pre></div><p>I&rsquo;m using <a href="https://github.com/VundleVim/Vundle.vim">Vundle</a> for my package management. Frankly, I should be using vim-plug, but I haven&rsquo;t taken the 5 minutes to move over. Vundle does the job just fine.</p>
<p>The first important package to have is <a href="https://github.com/fatih/vim-go">vim-go</a>. This plugin is absolutely fantastic, providing pretty much all the base go features you need. In my configuration for this plugin, I turn on syntax highlighting for basically everything. I also like <code>goimports</code> over <code>gofmt</code> because it sorts my imports, so I use that as the default formatter. Finally, I turn on highlighting of the variable my cursor is on, and I enable type hints in airline.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-vim" data-lang="vim"><span style="color:#75715e">&#34; use goimports not gofmt to format</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_fmt_command</span> = <span style="color:#e6db74">&#34;goimports&#34;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#75715e">&#34; syntax highlight all the things</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_build_constraints</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_extra_types</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_fields</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_functions</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_methods</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_operators</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_structs</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_highlight_types</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#75715e">&#34; highlight variables across file</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_auto_sameids</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#75715e">&#34; vim-go get type info in airline</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">go_auto_type_info</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span></code></pre></div><p>Once this is all set, and the plugin is installed, run <code>:GoInstallBinaries</code> to get the latest version of all required Go binaries.</p>
<h2 id="file-explorer">File Explorer</h2>
<p>One of the main reasons I switched from Neovim to VS Code for Go development in the first place was the presence of a persistent file browser. I had used NERDTree before, but the fact that it didn&rsquo;t persist through tabs made pretty useless for me. However, recently, I found out about <code>:NERDTreeMirror</code>, and it took NERDTree from a nice idea to a killer plugin for me.</p>
<p><code>:NERDTreeMirror</code> does exactly what it sounds like. It starts the NERDTree file browser in the current tab, mirroring the state of the other already-open trees. This allows NERDTree to effectively act like a persistent file browser. The only problem is that this command isn&rsquo;t automatically invoked when a new tab is opened, breaking the illusion of a persistent file browser. To fix this, I wrote my own new tab function that checks if NERDTree is already open, and if it is, <code>:NERDTreeMirror</code> is called after the new tab is created. I map this to <code>&lt;leader&gt;t</code>, which in my case, is <code>\t</code>.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-vim" data-lang="vim"><span style="color:#66d9ef">function</span>! <span style="color:#a6e22e">IsNerdTreeEnabled</span>()<span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>  <span style="color:#a6e22e">return</span> <span style="color:#a6e22e">exists</span>(<span style="color:#e6db74">&#39;t:NERDTreeBufName&#39;</span>) &amp;&amp; <span style="color:#a6e22e">bufwinnr</span>(<span style="color:#a6e22e">t</span>:<span style="color:#a6e22e">NERDTreeBufName</span>) != <span style="color:#ae81ff">-1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">endfunction</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">function</span>! <span style="color:#a6e22e">TreeTab</span>()<span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>  <span style="color:#66d9ef">if</span> <span style="color:#a6e22e">IsNerdTreeEnabled</span>()<span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>    <span style="color:#a6e22e">execute</span> <span style="color:#e6db74">&#39;tabe&#39;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>    <span style="color:#a6e22e">execute</span> <span style="color:#e6db74">&#39;NERDTreeMirror&#39;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>  <span style="color:#66d9ef">else</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>    <span style="color:#a6e22e">execute</span> <span style="color:#e6db74">&#39;tabe&#39;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span>  <span style="color:#66d9ef">endif</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">endfu</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">nnoremap</span> &lt;<span style="color:#a6e22e">leader</span>&gt;<span style="color:#a6e22e">t</span> :<span style="color:#a6e22e">call</span> <span style="color:#a6e22e">TreeTab</span>()&lt;<span style="color:#a6e22e">CR</span>&gt;<span style="color:#960050;background-color:#1e0010">
</span></code></pre></div><p>My other config settings, in order</p>
<ol>
<li>Automatically open NERDTree if no arguments are passed into <code>vim</code></li>
<li>Open NERDTree to the current file when <code>\v</code> is typed</li>
<li>Close NERDTree if it is the only window left</li>
</ol>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-vim" data-lang="vim"><span style="color:#75715e">&#34; auto-open if no args are set</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">autocmd</span> <span style="color:#a6e22e">VimEnter</span> * <span style="color:#66d9ef">if</span> !<span style="color:#a6e22e">argc</span>() | <span style="color:#a6e22e">NERDTree</span> | <span style="color:#66d9ef">endif</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#75715e">&#34; open on \v</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">nnoremap</span> &lt;<span style="color:#a6e22e">silent</span>&gt; &lt;<span style="color:#a6e22e">Leader</span>&gt;<span style="color:#a6e22e">v</span> :<span style="color:#a6e22e">NERDTreeFind</span>&lt;<span style="color:#a6e22e">CR</span>&gt;<span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#75715e">&#34; close if only window left</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">autocmd</span> <span style="color:#a6e22e">bufenter</span> * <span style="color:#66d9ef">if</span> (<span style="color:#a6e22e">winnr</span>(<span style="color:#e6db74">&#34;$&#34;</span>) == <span style="color:#ae81ff">1</span> &amp;&amp; <span style="color:#a6e22e">exists</span>(<span style="color:#e6db74">&#34;b:NERDTree&#34;</span>) &amp;&amp; <span style="color:#a6e22e">b</span>:<span style="color:#a6e22e">NERDTree</span>.<span style="color:#a6e22e">isTabTree</span>()) | <span style="color:#a6e22e">q</span> | <span style="color:#66d9ef">endif</span><span style="color:#960050;background-color:#1e0010">
</span></code></pre></div><h2 id="linting">Linting</h2>
<p>I use <a href="https://github.com/w0rp/ale">Ale</a> for linting. It&rsquo;s very fast, and most importantly, it&rsquo;s dumb easy to set up. I&rsquo;ve struggled for hours with Neomake only to get it half-working, but Ale worked out of the box. Just make sure you have <a href="https://github.com/golang/lint">golint</a> installed, and you&rsquo;re good to go.</p>
<p>I also manually configure the error and warning characters and disable the location list for errors. It&rsquo;s more annoying for me than it is helpful.</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-vim" data-lang="vim"><span style="color:#75715e">&#34; Error and warning signs.</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">ale_sign_error</span> = <span style="color:#e6db74">&#39;⤫&#39;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">ale_sign_warning</span> = <span style="color:#e6db74">&#39;⚠&#39;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">ale_set_loclist</span> = <span style="color:#ae81ff">0</span><span style="color:#960050;background-color:#1e0010">
</span></code></pre></div><h2 id="autocomplete">Autocomplete</h2>
<p>Autocomplete is notoriously hard to set up in Vim/Neovim, and unfortunately, I don&rsquo;t have silver bullet to fix that. I&rsquo;m using <a href="https://github.com/Shougo/deoplete.nvim">deoplete</a> and <a href="https://github.com/deoplete-plugins/deoplete-go">deoplete-go</a> for better Go support. To use the latter, make sure you have <a href="https://github.com/mdempsky/gocode">gocode</a> installed. I&rsquo;m still having some problems with errors on startup, but the solution is very good once <code>gocode</code> has a chance to start up.</p>
<p>My config options, in order</p>
<ol>
<li>start deoplete at startup</li>
<li>allow me to cycle through autocomplete options with the tab key</li>
<li>set a order of preference for go autocomplete suggestions</li>
</ol>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-vim" data-lang="vim"><span style="color:#75715e">&#34; deoplete settings</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">deoplete</span>#<span style="color:#a6e22e">enable_at_startup</span> = <span style="color:#ae81ff">1</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#a6e22e">inoremap</span> &lt;<span style="color:#a6e22e">expr</span>&gt;&lt;<span style="color:#a6e22e">TAB</span>&gt;  <span style="color:#a6e22e">pumvisible</span>() ? <span style="color:#e6db74">&#34;\&lt;C-n&gt;&#34;</span> : <span style="color:#e6db74">&#34;\&lt;TAB&gt;&#34;</span><span style="color:#960050;background-color:#1e0010">
</span><span style="color:#960050;background-color:#1e0010"></span><span style="color:#66d9ef">let</span> <span style="color:#a6e22e">g</span>:<span style="color:#a6e22e">deoplete</span>#<span style="color:#a6e22e">sources</span>#<span style="color:#a6e22e">go</span>#<span style="color:#a6e22e">sort_class</span> = [<span style="color:#e6db74">&#39;package&#39;</span>, <span style="color:#e6db74">&#39;func&#39;</span>, <span style="color:#e6db74">&#39;type&#39;</span>, <span style="color:#e6db74">&#39;var&#39;</span>, <span style="color:#e6db74">&#39;const&#39;</span>]<span style="color:#960050;background-color:#1e0010">
</span></code></pre></div><p>A few tips for setup:</p>
<ul>
<li>
<p>Make sure the python and python3 Neovim providers are installed. They can be installed with the following commands:</p>
<pre><code>$ pip install neovim
$ pip3 install neovim
</code></pre></li>
<li>
<p>Make sure all necessary Go binaries are installed. Ensure this with the <code>:GoInstallBinaries</code> command in Neovim.</p>
</li>
<li>
<p>Add a make instruction in Vundle for <code>deoplete-go</code> so that it gets made properly.</p>
<pre><code>Plugin 'zchee/deoplete-go', { 'do': 'make'}
</code></pre></li>
</ul>
<h2 id="misc">Misc</h2>
<p>Some other plugins I use are:</p>
<ul>
<li><a href="https://github.com/jiangmiao/auto-pairs">auto-pairs</a> for automatically closing parentheses and brackets</li>
<li><a href="https://github.com/vim-airline/vim-airline">vim-airline</a> for a fantastic statusline</li>
<li><a href="https://github.com/ryanoasis/vim-devicons">vim-devicons</a> to add type icons to NERDTree</li>
<li><a href="https://github.com/ctrlpvim/ctrlp.vim">ctrlp.vim</a> for fast fuzzy search</li>
</ul>
<h2 id="conclusion">Conclusion</h2>
<p>With these plugins, I can now use Neovim without losing any of the functionality I got in VS Code! If you&rsquo;re interested, <a href="https://github.com/Pwpon500/user-sync/blob/master/vimrc">here&rsquo;s my vimrc</a>. Happy coding!</p>
]]></content></item><item><title>Caplance Development Update 1</title><link>https://pawa.lt/posts/2019/05/caplance-development-update-1/</link><pubDate>Sat, 25 May 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/05/caplance-development-update-1/</guid><description>After countless hours spent reading the GoDocs, googling &amp;ldquo;gre packets golang&amp;rdquo;, and staring at the term SIGSEV, I&amp;rsquo;ve got an update! I&amp;rsquo;ve made pretty significant progress on the actual load balancer, implementing most of the base functionality. So far, the load balancer can:
listen for and ingest whole IP packets select the backend for a packet based on a combination of IP and source port encapsulate packets in UDP and send them to the appropriate backend attach a specified virtual IP I&amp;rsquo;ve also implemented some extremely rudimentary testing and setup, but I&amp;rsquo;m going to rework those once the project structure solidifies some more.</description><content type="html"><![CDATA[<p>After countless hours spent reading the GoDocs, googling &ldquo;gre packets golang&rdquo;, and staring at the term SIGSEV, I&rsquo;ve got an update! I&rsquo;ve made pretty significant progress on the actual load balancer, implementing most of the base functionality. So far, the load balancer can:</p>
<ul>
<li>listen for and ingest whole IP packets</li>
<li>select the backend for a packet based on a combination of IP and source port</li>
<li>encapsulate packets in UDP and send them to the appropriate backend</li>
<li>attach a specified virtual IP</li>
</ul>
<p>I&rsquo;ve also implemented some extremely rudimentary testing and setup, but I&rsquo;m going to rework those once the project structure solidifies some more.</p>
<p>Here&rsquo;s some detail on the more interesting technical parts of the project so far:</p>
<h2 id="listening">Listening</h2>
<p>I initially listened for packets with a go interface for pcap, a packet capture protocol that allows the user to select an interface and &ldquo;sniff&rdquo; its packets. If you&rsquo;ve used tcpdump, you&rsquo;ve used pcap. With pcap, I could tap into the interface that the virtual IP was using and filter for only the packets destined for the VIP. Then, I could read the packets right off the interface. There were two problems with this approach:</p>
<ol>
<li>I wasn&rsquo;t actually doing listening in a conventional way. If you pull up <code>ss -4l</code>, you wouldn&rsquo;t see any entry for Caplance (because I&rsquo;m not technically listening. I&rsquo;m creating a TAP device). That&rsquo;s a pretty big problem just because I want Caplance to work well with conventional Unix tools.</li>
<li>Closing a pcap device from go is stupid slow. In my tests, closing the device took upwards of 30 seconds, and the device needs to be closed before Caplance can terminate. This is unacceptable, and it was ultimately what did in the pcap approach for me.</li>
</ol>
<p>Enter <a href="https://golang.org/pkg/net/#IPConn">IPConn</a>. IPConn is a listener for go that listens at the IP level. That means, when you do a read off of it, it reads in from the IP layer down, exactly what I want. Since it&rsquo;s a built-in go listener, it also closes extremely quickly. Whereas before I needed to set up a filter, find the appropriate device, etc., with IPConn, my listening can be reduced to this simple function:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> (<span style="color:#a6e22e">b</span> <span style="color:#f92672">*</span><span style="color:#a6e22e">Balancer</span>) <span style="color:#a6e22e">listenWithConn</span>(<span style="color:#a6e22e">conn</span> <span style="color:#f92672">*</span><span style="color:#a6e22e">net</span>.<span style="color:#a6e22e">IPConn</span>, <span style="color:#a6e22e">pool</span> <span style="color:#f92672">*</span><span style="color:#a6e22e">sync</span>.<span style="color:#a6e22e">Pool</span>) {
	<span style="color:#66d9ef">for</span> {
		<span style="color:#a6e22e">buf</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">pool</span>.<span style="color:#a6e22e">Get</span>().([]<span style="color:#66d9ef">byte</span>)
		<span style="color:#a6e22e">n</span>, <span style="color:#a6e22e">err</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">conn</span>.<span style="color:#a6e22e">Read</span>(<span style="color:#a6e22e">buf</span>)
		<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">err</span> <span style="color:#f92672">!=</span> <span style="color:#66d9ef">nil</span> {
			<span style="color:#a6e22e">log</span>.<span style="color:#a6e22e">Println</span>(<span style="color:#e6db74">&#34;could not read from connection&#34;</span>)
			<span style="color:#66d9ef">continue</span>
		}
		<span style="color:#a6e22e">toSend</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">rawPacket</span>{<span style="color:#a6e22e">buf</span>, <span style="color:#a6e22e">n</span>}
		<span style="color:#a6e22e">b</span>.<span style="color:#a6e22e">packets</span> <span style="color:#f92672">&lt;-</span> <span style="color:#a6e22e">toSend</span>
	}
}
</code></pre></div><p>To provide a bit of context around this function, <code>pool</code> is a pool of buffers that the IPConn can use to store data from reads. <code>b.packets</code> is a channel (concurrency-safe queue) of packet for the worker goroutines to read packets off of.</p>
<h2 id="packet-forwarding">Packet Forwarding</h2>
<p>While I would to use normal routing rules to route incoming packets over GRE tunnels to backends, that won&rsquo;t work here. If we were always forwarding the same IP range of packets to the same backends, we could use iptables, but for this project, we want to use consistent hashing. Because of this, we have to write listening (as seen above) and packet forwarding ourselves.</p>
<p>First, it&rsquo;s important to note that there&rsquo;s no good GRE library for go right now. The only real way to create a GRE tunnel in go is to make calls to the kernel to tell it to construct one. There&rsquo;s no way to have a GRE object, for instance, that you can just write data to in the same way you do with a TCP connection. I may take this up as a future project.</p>
<p>Because of the lack of a library, I would have to write data directly to the wire. My initial idea was to (again) use pcap. My algorithm was as follows:</p>
<ol>
<li>Create GRE tunnel to backend (with netlink)</li>
<li>Manually create GRE header for backend</li>
<li>For each packet destined for the backend:
<ol>
<li>Create appropriate GRE header for the packet</li>
<li>Marshal GRE header into a <code>[]byte</code></li>
<li>Append the packet onto the marshalled header</li>
<li>Perform some crazy pcap magic to figure out mac addresses of the next hop on the way to the backend</li>
<li>Write the encapsulated packet to the right physical interface</li>
</ol>
</li>
</ol>
<p>Ya know, now that I type it all out, that was an awful plan.</p>
<p>After a lot of frustration and reflection, I asked myself the pivotal question, &ldquo;Why am I so set on GRE in the first place?&rdquo; The only reason I could come up with was that Google used it, and that&rsquo;s a pretty bad reason. Thinking about it more, I realized I could just use a UDP connection in place of GRE and achieve the same results I wanted.</p>
<p>Now, my algorithm is as follows:</p>
<ol>
<li>Dial a UDP connection to each backend</li>
<li>Write each packet to the appropriate backend with the built-in <code>Write(data []byte)</code> method</li>
</ol>
<p>That&rsquo;s a wee bit simpler. It&rsquo;s not as easy to show in a single piece of code as listening, but it wasn&rsquo;t too hard to write.</p>
<h2 id="graceful-stop">Graceful Stop</h2>
<p>I implemented graceful stop, which turned out to be way harder than I expected. I&rsquo;m anticipating writing about it in the future, though, so I&rsquo;m not going to include details on how it works right now. It&rsquo;s also very likely that I&rsquo;ll change how it&rsquo;s implemented. If you want to see the current iteration of it right now, <a href="https://github.com/Pwpon500/caplance/blob/4887f8c6230fbe062660c300df0a81f02450f064/balancer/control.go#L80">you can find the code here</a>.</p>
<h1 id="future-work">Future Work</h1>
<p>Obviously, I&rsquo;m not done yet. I still need to implement the following:</p>
<ul>
<li>Configuration file</li>
<li>Environment variable configuration</li>
<li>Client registration</li>
<li>Client health checks</li>
<li>Client packet injestion</li>
<li>Optimization</li>
<li>More that I don&rsquo;t even know exists yet</li>
</ul>
<p>I&rsquo;m crazy excited to get started on all those items, and I can&rsquo;t wait to write about them once I figure them out.</p>
]]></content></item><item><title>Constant Improvement - A Reflection</title><link>https://pawa.lt/posts/2019/05/constant-improvement-a-reflection/</link><pubDate>Tue, 21 May 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/05/constant-improvement-a-reflection/</guid><description>From Sunday, 3/10/19 to today, Tuesday 5/21/19, I&amp;rsquo;ve been working on improving. During that time, I decided on a single, small improvement to make each week, and I tried to execute on it. Each week, I&amp;rsquo;d outline the motivation behind my decision, the tools I&amp;rsquo;d be using, and the rules I&amp;rsquo;d have to follow during the week. The next week, I&amp;rsquo;d write a short reflection on how I thought the week went with respect to my adherence to the rules I set.</description><content type="html"><![CDATA[<p>From Sunday, 3/10/19 to today, Tuesday 5/21/19, <a href="https://github.com/Pwpon500/improvement">I&rsquo;ve been working on improving</a>. During that time, I decided on a single, small improvement to make each week, and I tried to execute on it. Each week, I&rsquo;d outline the motivation behind my decision, the tools I&rsquo;d be using, and the rules I&rsquo;d have to follow during the week. The next week, I&rsquo;d write a short reflection on how I thought the week went with respect to my adherence to the rules I set.</p>
<p>During this time, I learned a significant amount about myself and the process of improvement, so I thought I&rsquo;d share a bit about what I&rsquo;ve learned and some of my personal opinions on the topic. While reading, keep in mind that what I&rsquo;m saying is based purely off my own experience and is likely not applicable to all who read this.</p>
<h1 id="making-improvements-small">Making improvements small</h1>
<p>I&rsquo;ve made the mistake of trying to make large improvements many times. The danger with this is that you set expectations that you will almost certainly not be able to meet. The subsequent dropping of some of the expectations just weakens the others, and eventually, the whole system falls apart. The hard part is that the desire to make huge improvements often comes from the best place in the heart! Wanting to improve is an admirable thing, and wanting to radically improve yourself is even more admirable. Unfortunately, for the vast majority of people, I just don&rsquo;t think quick radical improvement is feasible.</p>
<p>Obviously, the solution here is to make your improvements small. For me, the easiest way to make my improvements small is to first make the rules for my improvement unambiguously specific. After the rules have been fully specified, I can easily see if what I&rsquo;m doing is too much or not. I work on those rules for a week, and afterward, I reflect on whether or not they were too strenuous. Typically, I can tell that they were too strenuous simply if I was able to stick ot them or not.</p>
<h1 id="allowing-rules-to-change">Allowing rules to change</h1>
<p>You&rsquo;re never going to get your rules right the first time. In fact, they shouldn&rsquo;t! That&rsquo;s why you try out the rules in 1-week increments in the first place. The important thing is to track what rules are useful, which aren&rsquo;t, and what rules you add/change throughout the week.</p>
<p>For example, in my week 1, I decided to track my time with an app called Clockify. The only problem is that Clockify didn&rsquo;t help me at all. Timing myself on tasks made near no difference. However, halfway through the week, I found an app called Forest that was hugely helpful for me. It not only timed me but also locked me out of my phone, which ended up being the kick I needed. I noted that I was using Forest instead of Clockify in my reflection, and the change was done.</p>
<p>I should note, however, that allowing rules to change doesn&rsquo;t permit just dropping rules. If you just start dropping rules, it cheapens the values of the rest, and eventually, you&rsquo;ll give up on those. If you are getting rid of a rule or changing a rule, it&rsquo;s imperative that you have a good reason for doing so and that you record the change.</p>
<h1 id="the-eventual-end-of-it-all">The eventual end of it all</h1>
<p>At the end of the day, making an improvement a week, no matter how small, piles up (that&rsquo;s the idea). For me, after about 7 weeks of the project, I was done. The project had given me a lot to think about, and it really helped me get on track at a time where my life was pretty disorderly.</p>
<p>While I&rsquo;m done adding improvements for the time being, I&rsquo;m still using many of the improvements I made during the project! I&rsquo;m still:</p>
<ul>
<li>Tracking tasks in Todoist</li>
<li>Focusing with Forest</li>
<li>Exercising 5+ times/week</li>
<li>Paying attention to my screen time and reading my reports</li>
<li>Blogging (obviously)</li>
<li>Budgeting with Mint</li>
</ul>
<p>I&rsquo;d really encourage anyone reading to try out a project like this. It&rsquo;s certainly helped me get myself on track in ways that I&rsquo;ve been attempting for years now. There&rsquo;s a link to my Github repo at the top, but I&rsquo;ll drop it again here:</p>
<p><a href="https://github.com/Pwpon500/improvement">https://github.com/Pwpon500/improvement</a></p>
]]></content></item><item><title>Mininet for Education</title><link>https://pawa.lt/posts/2019/04/mininet-for-education/</link><pubDate>Mon, 29 Apr 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/04/mininet-for-education/</guid><description>For about 2 years, I built and led the NEAT Rack Project, a program meant to teach high school and college aged students network engineering. In this program, we covered things as basic as introduction to Linux and as advanced as firewalling or dynamic routing with BGP. This program was well-run, well-written, and it overall did a solid job of giving students an introduction to technology infrastructure. There was just one problem with it - hardware.</description><content type="html"><![CDATA[<p>For about 2 years, I built and led the <a href="http://rva-ix.net/the-neat-rack-program/">NEAT Rack Project</a>, a program meant to teach high school and college aged students network engineering. In this program, we covered things as basic as introduction to Linux and as advanced as firewalling or dynamic routing with BGP. This program was well-run, well-written, and it overall did a solid job of giving students an introduction to technology infrastructure. There was just one problem with it - hardware.</p>
<p>This program requires that the students have access to a rack with a managed switch and at least one server capable of virtualization. First, these resources are often difficult for schools to procure. Then, even if the school can procure them, only a single student or group of students can work on the labs at a time. Furthermore, working on labs at home isn&rsquo;t even a possibility since you can&rsquo;t take a whole rack home.</p>
<p>Mininet aims to solve these problems.</p>
<h2 id="what-is-mininet">What is Mininet?</h2>
<p>Mininet is a network simulation tool for Linux. The goal of it is to create a &ldquo;realistic virtual network&rdquo; with minimal overhead. Since it aims for minimal overhead, it shares both filesystem space and PID space with the host it operates on. Both of these can be avoided, however, with the private directories host option in Mininet.</p>
<p>One cool thing about Mininet is how lightweight it is. Since it does basically no isolation other than creating some virtual kernels, it spins up and down in seconds, using minimal memory. This means it can run on a device as weak as a Raspberry Pi!</p>
<h2 id="the-mininet-api">The Mininet API</h2>
<p>Another amazing part of Mininet is its Python API. This API lets you programmatically create new networks, interact with them, and destroy them. Take the following simple example:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-python" data-lang="python"><span style="color:#75715e">#!/usr/bin/env python2</span>
<span style="color:#e6db74">&#34;&#34;&#34; This script provides a basic switch topo.&#34;&#34;&#34;</span>

<span style="color:#f92672">from</span> functools <span style="color:#f92672">import</span> partial
<span style="color:#f92672">from</span> mininet.topo <span style="color:#f92672">import</span> SingleSwitchTopo
<span style="color:#f92672">from</span> mininet.net <span style="color:#f92672">import</span> Mininet
<span style="color:#f92672">from</span> mininet.cli <span style="color:#f92672">import</span> CLI
<span style="color:#f92672">from</span> mininet.node <span style="color:#f92672">import</span> Host
<span style="color:#f92672">from</span> mininet.log <span style="color:#f92672">import</span> setLogLevel


<span style="color:#66d9ef">def</span> <span style="color:#a6e22e">simple_test</span>():
    <span style="color:#e6db74">&#34;Create and test a simple network&#34;</span>
    topo <span style="color:#f92672">=</span> SingleSwitchTopo(k<span style="color:#f92672">=</span><span style="color:#ae81ff">5</span>)
    private_dirs <span style="color:#f92672">=</span> [<span style="color:#e6db74">&#39;/run&#39;</span>, (<span style="color:#e6db74">&#39;/var/run&#39;</span>, <span style="color:#e6db74">&#39;/tmp/</span><span style="color:#e6db74">%(name)s</span><span style="color:#e6db74">/var/run&#39;</span>), <span style="color:#e6db74">&#39;/var/mn&#39;</span>]
    host <span style="color:#f92672">=</span> partial(Host, privateDirs<span style="color:#f92672">=</span>private_dirs)
    net <span style="color:#f92672">=</span> Mininet(topo<span style="color:#f92672">=</span>topo, host<span style="color:#f92672">=</span>host)
    net<span style="color:#f92672">.</span>start()
    CLI(net)
    net<span style="color:#f92672">.</span>stop()


<span style="color:#66d9ef">if</span> __name__ <span style="color:#f92672">==</span> <span style="color:#e6db74">&#39;__main__&#39;</span>:
    <span style="color:#75715e"># Tell mininet to print useful information</span>
    setLogLevel(<span style="color:#e6db74">&#39;info&#39;</span>)
    simple_test()
</code></pre></div><p>In these few lines of code, I have a file that will automatically create a network with 5 hosts on it where each host has PID separation. Now, instead of having to manually specify all these options, the user can just run &ldquo;sudo ./setup.py&rdquo;, and they&rsquo;re consoled in.</p>
<p>The possibilities only begin here. Things get even crazier with <a href="http://mininet.org/walkthrough/#custom-topologies">custom topologies</a>.</p>
<h2 id="teaching-example">Teaching Example</h2>
<p>I could go on and on about Mininet, but the power of it becomes most evident when you actually see how it can be used. <a href="https://github.com/pwpon500/teaching">Here</a>, I&rsquo;m working on building up some labs based around Mininet for my friends to learn network engineering with. Feel free to try the labs and contribute your own!</p>
]]></content></item><item><title>Maglev - A Next-Generation Load Balancer</title><link>https://pawa.lt/posts/2019/04/maglev-a-next-generation-load-balancer/</link><pubDate>Sat, 20 Apr 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/04/maglev-a-next-generation-load-balancer/</guid><description>When reading Google&amp;rsquo;s SRE book, I came across the section on load balancing. The book breaks up load balancing into three levels:
DNS level Network level/Virtual IP level Datacenter level While I was familiar with the Datacenter level (aka reverse proxy load balancing) and DNS load balancing (achieved partially through round robin DNS), I had never looked at network-level load balancing. After reading this section, I was fascinated with how their network-level load balancing works.</description><content type="html"><![CDATA[<p>When reading <a href="https://landing.google.com/sre/sre-book/toc/index.html">Google&rsquo;s SRE book</a>, I came across the section on load balancing. The book breaks up load balancing into three levels:</p>
<ol>
<li>DNS level</li>
<li>Network level/Virtual IP level</li>
<li>Datacenter level</li>
</ol>
<p>While I was familiar with the Datacenter level (aka reverse proxy load balancing) and DNS load balancing (achieved partially through round robin DNS), I had never looked at network-level load balancing. After reading this section, I was fascinated with how their network-level load balancing works.</p>
<p>After some research online, I realized that the system they were referring to was their load balancer called Maglev. In 2016, Google released a paper detailing Maglev and how they built it. If you&rsquo;re interested in the nitty-gritty details, you can read the paper <a href="https://ai.google/research/pubs/pub44824">here</a>. It&rsquo;s a dense but fascinating read. I could spend a huge amount of time going into all the details of what makes Maglev interesting, but instead, I&rsquo;ll focus on two aspects: Maglev hashing and packet encapsulation.</p>
<h2 id="introduction">Introduction</h2>
<p>Before anything else, I&rsquo;ll describe the function of a network load balancer. The job of a network load balancer is to distribute packets destined for a &ldquo;virtual IP address&rdquo; to a set of backends evenly. Take the following picture as an example:</p>
<p><img src="/img/maglev_1.png" alt="example maglev"></p>
<p>In this example, the job of Maglev is to deliver all of the packets coming in from the public internet evenly to backends 1-4.</p>
<h2 id="maglev-hashing">Maglev Hashing</h2>
<p>The first part of load balancing is actually deciding which backend to send an incoming packet to. In network load balancing, we are just looking at the IP layer, so we can&rsquo;t keep track of things like which backend has the most TCP connections open to it. Since we can&rsquo;t keep track of current state, we have to figure out how to distribute packets evenly just based on some attribute of the packet. Assume we pick out some attribute of each packet (this is usually a hashed combination of source IP and source port) and call id <code>id</code>. Then, we can treat <code>id</code> as an integer and evenly distribute packets to backends using the function <code>backend = id (mod n)</code> where <code>n</code> is our number of backends. Assuming an even distribution of IDs, this method will evenly distribute packets to backends, and it will keep sending packets from the same source to the same backend.</p>
<p>So, is it that simple? It seems like we&rsquo;ve achieved what we want to - even distribution - with no downsides. Consider what happens when our <code>n</code> increases or decreases by even 1, however. In this case, since we don&rsquo;t know the range of our <code>id</code> hash function, which backend our packets go to could be completely changed for all connections. This would mean that all open client connections would be broken, which is certainly not desirable behavior. To see this in action, let&rsquo;s look at the example of where to send incoming packets with 5 backends versus with 4 backends.</p>
<table>
<thead>
<tr>
<th>id</th>
<th>id (mod 4)</th>
<th>id (mod 5)</th>
</tr>
</thead>
<tbody>
<tr>
<td>717</td>
<td>1</td>
<td>2</td>
</tr>
<tr>
<td>561</td>
<td>1</td>
<td>1</td>
</tr>
<tr>
<td>544</td>
<td>0</td>
<td>4</td>
</tr>
<tr>
<td>67</td>
<td>3</td>
<td>2</td>
</tr>
<tr>
<td>310</td>
<td>2</td>
<td>0</td>
</tr>
<tr>
<td>626</td>
<td>2</td>
<td>1</td>
</tr>
</tbody>
</table>
<p>As you can see, all the packet flows are disrupted except that for the packet with id 561. Since we will certainly be adding and removing backends live, this is unacceptable behavior. This is where <strong>consistent hashing</strong> comes to the rescue. Consistent hashing is very complicated, but the short of it is that it provides a way for minimal disruption when adding and removing backends. With consistent hashing, instead of all the packet flows being disrupted, at most <code>1/n</code> (where <code>n</code> is the number of the backends) of the packets flows are disrupted.</p>
<p>There are many trade-offs with different kinds of consistent hashing that <a href="https://medium.com/@dgryski/consistent-hashing-algorithmic-tradeoffs-ef6b8e2fcae8">this article</a> explains better than I can, but in essence, there are three things you can optimize for in consistent hashing: memory usage, lookup speed, and hashtable rebuild speed. You can typically have 2 of these at once, but you can never have all 3 without making some other very significant tradeoff. Google&rsquo;s solution, called <strong>Maglev Hashing</strong>, optimizes for the first two. It assumes node failures are uncommon, and in making that assumption, it can get low memory usage with high lookup speed and minimal disruption when <em>new</em> backends are added but poor rebuild speed when a backend fails. This is a pretty reasonable assumption, and it turns out to work extremely well in the context of network load balancing.</p>
<p>Now that we know how to pick backends, let&rsquo;s actually talk about how we send data to backends.</p>
<h2 id="packet-encapsulation">Packet Encapsulation</h2>
<p>First, let&rsquo;s consider the reverse proxy method of load balancing. In this method, connections are taken in by the load balancer. The balancer then initiates a new connection to the appropriate backend for the connection, and proxies the connections together, allowing the client to indirectly talk to a backend. This method looks like this:</p>
<p><img src="/img/maglev_2.png" alt="reverse proxy design"></p>
<p>In this design, the load balancer has to do the following:</p>
<ul>
<li>Keep track of all active connections</li>
<li>Proxy data between connections for all connections</li>
<li>Pass both ingress and egress traffic for all the backends</li>
</ul>
<p>While this is acceptable (and even desired) for a datacenter load balancer, we can&rsquo;t pass packets at the scale a network load balancer needs to using this design. This is where <strong>packet encapsulation</strong> comes into play.</p>
<p>The solution that Google found to this problem was to use the following algorithm:</p>
<ol>
<li>Maglev hash the packet to determine its backend</li>
<li>Wrap the packet in a layer of GRE (generic routing encapsulation)</li>
<li>Send the encapsulated packet to the desired backend</li>
<li>Have the backend break the packet of the encapsulation and fully process it</li>
<li>Have the backend directly reply to this packet.</li>
</ol>
<p>Did you catch that last step? Since we&rsquo;re using encapsulation, the backend is seeing the original packet as received by the load balancer, so it can directly reply to the packet. We can see this in action on an example packet flow:</p>
<p><img src="/img/maglev_3.png" alt="gre design"></p>
<p>The fact that the backend can directly reply is a huge benefit of the encapsulation method. To see why, consider the use case of YouTube. When a user requests to see a video, the request payload is very small, only containing metadata about what video they want to watch. The reply, however, is gigantic since it&rsquo;s a full video of up to 8K quality! With this method, the load balancer only only has to worry about the request, meaning it can handle drastically more traffic than a reverse proxy can.</p>
<p>Personally, I thought this was the most interesting part of Maglev. I had only ever thought to use GRE as a site-to-site VPN, but this gave me a whole new outlook on what it could do.</p>
<h2 id="conclusion">Conclusion</h2>
<p>Maglev brings a ton of interesting concepts to the table and combines them in a way we&rsquo;ve never really seen before. I could only cover two things here, but tech like ECMP and RPC also play a huge role in how Maglev works. If you&rsquo;re interested, I highly encourage reading the paper.</p>
<p>Thanks for reading! Look below for a shameless plug.</p>
<h2 id="shameless-plug">Shameless Plug</h2>
<p>As I said, I found Google&rsquo;s ideas on network load balancers fascinating. In fact, I found it so interesting that I decided to implement it myself! If you&rsquo;re interested in seeing how some of these ideas are implemented, I&rsquo;m writing my own version of Maglev called Caplance <a href="https://github.com/pwpon500/caplance">here</a>. At the time of writing this, I&rsquo;m working on setting up automated testing with Mininet, and I&rsquo;ll probably write more in the coming weeks as I make more progress on the load balancer.</p>
]]></content></item><item><title>On the Value of Formal Education</title><link>https://pawa.lt/posts/2019/04/on-the-value-of-formal-education/</link><pubDate>Sun, 14 Apr 2019 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2019/04/on-the-value-of-formal-education/</guid><description>Over the past few years, I&amp;rsquo;ve started to see a multitude of online videos/articles about the weakening value of college, particularly in the area of computer science. Whether it&amp;rsquo;s Devon Crawford talking about how he dropped out of school or Mashable writing an article trashing CS degrees, the popular view on the internet seems to be &amp;ldquo;Eh, just learn it on your own.&amp;rdquo; I&amp;rsquo;ve struggled with my opinion significantly on this topic, but overall, I think this view is deeply flawed.</description><content type="html"><![CDATA[<p>Over the past few years, I&rsquo;ve started to see a multitude of online videos/articles about the weakening value of college, particularly in the area of computer science. Whether it&rsquo;s <a href="https://www.youtube.com/channel/UCDrekHmOnkptxq3gUU0IyfA">Devon Crawford</a> talking about how he dropped out of school or <a href="https://mashable.com/2014/12/16/warning-college-may-be-a-waste-of-your-time-and-money/#5MchVc1AFaqc">Mashable writing an article trashing CS degrees</a>, the popular view on the internet seems to be &ldquo;Eh, just learn it on your own.&rdquo; I&rsquo;ve struggled with my opinion significantly on this topic, but overall, I think this view is deeply flawed. First, though, I&rsquo;ll talk about where this view goes right.</p>
<h2 id="who-needs-college-really">Who needs college really?</h2>
<p>The argument against a CS degree typically takes two main points:</p>
<ul>
<li>You can learn CS on your own</li>
<li>Any current technologies won&rsquo;t be taught in school anyway</li>
</ul>
<p>I think the first point is much less valid than people give it credit for. It&rsquo;s easy to learn <em>programming</em> on your own, but learning computer science on your own is far more challenging. While programming is perhaps the most critical part of computer science, being a good programmer is a drastically different thing from being a good software engineer. Concepts like proofs, time complexity, and complex algorithms are also crucial to fully developing as a software engineer, and these concepts are extremely challenging to learn on your own. I&rsquo;ll go more into this idea in the next section.</p>
<p>This next point is actually where I think the aforementioned Devon Crawford gets it right. <strong>Colleges simply cannot stay current with state-of-the-art technologies.</strong> Take <a href="https://www.youtube.com/watch?v=SC7lLm6QAb8">Devon&rsquo;s example of Kubernetes</a>. Kubernetes is the state-of-the-art in container orchestration, and it has pretty much taken over the container orchestration space, at least in open source. Despite all of this, most universities don&rsquo;t have classes even touching it, much less actually going into how it works. This is a huge weakness for colleges simply because of how much hiring is skills-based. Having something like &ldquo;Kubernetes proficiency&rdquo; is a huge boost to any CS resume, but because of the nature of how long it takes to build and establish a new class, colleges will never be able to keep up.</p>
<h2 id="the-case-for-college">The case for college</h2>
<p>To touch on the self-teaching side of CS, let me provide a personal anecdote:</p>
<p>In my junior year of high school, my programming competition team started to become seriously competitive. We started to place highly at tournaments, and it became clear that if we wanted to reach the next level, we would have to learn real algorithms. It became my job to learn all things graph theory. In the span of a few weeks, I learned BFS, DFS, Dijkstra&rsquo;s, and Prim&rsquo;s - or so I thought.</p>
<p>Fast forward to my second semester of college. Now, I&rsquo;m re-learning all of those algorithms, and I&rsquo;m realizing just how wrong I was about all of them. For one thing, my old versions of the algorithms all ran in at least n squared time, if not n cubed. Getting a formal education around these has taught me how to do these in linear or logarithmic time. My understanding of them has also far deepened, teaching me how to use them in a variety of different contexts I would have never imagined before.</p>
<p>It&rsquo;s primarily this experience that has made me understand the value of a formal education. While I &ldquo;knew&rdquo; these algorithms before, I now realize how little I actually understood about them. I think idea holds for a multitude of computer science concepts.</p>
<p>To address colleges&rsquo; difficulty in teaching new technologies, I actually don&rsquo;t think that&rsquo;s the job of the college. <strong>The job of the college is to provide the student with the necessary background as to be able to learn new tools quickly.</strong> The job of the college is not to teach current technologies, because if it was, those skills would be useless when that tool fell out of favor. The best thing a college can do is what the good ones are doing - provide the students with a solid foundation to build on. If the student has the right background, they&rsquo;ll be able to learn any tool you can throw at them.</p>
<p>College also provides a multitude of other benefits not listed here including, but not limited to, establishing professional connections, facilitating learning outside one&rsquo;s comfort zone, and providing a safe space for young people to grow into the world.</p>
<h2 id="conclusion">Conclusion</h2>
<p>While I understand the ideas behind the nay-sayers of college, I ultimately disagree with them. I think that in the long-term, going to college is one of the best decisions you could make.</p>
]]></content></item><item><title>Hacking Farkle</title><link>https://pawa.lt/posts/2018/12/hacking-farkle/</link><pubDate>Sat, 29 Dec 2018 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2018/12/hacking-farkle/</guid><description>Recently, my family has been playing Farkle, a simple but difficult-to-analyze dice game. Farkle has the following rules:
At the start of the turn, each player must roll the 6 dice. After rolling, the player sets aside all the dice that scored. The player can then choose to roll again to potentially get more points or to take the points they have already scored. If the player rolls and scores no points, they get 0 points and their turn is over.</description><content type="html"><![CDATA[<p>Recently, my family has been playing Farkle, a simple but difficult-to-analyze dice game. Farkle has the following rules:</p>
<ul>
<li>At the start of the turn, each player must roll the 6 dice.</li>
<li>After rolling, the player sets aside all the dice that scored.</li>
<li>The player can then choose to roll again to potentially get more points or to take the points they have already scored.</li>
<li>If the player rolls and scores no points, they get 0 points and their turn is over.</li>
<li>If the player uses up all 6 dice, they can recycle all of them and roll all 6 again.</li>
</ul>
<p>The ways to score are as follows (dice cannot be double counted when scoring):</p>
<table>
<thead>
<tr>
<th>Combination</th>
<th>Points</th>
</tr>
</thead>
<tbody>
<tr>
<td>Single 1</td>
<td>100</td>
</tr>
<tr>
<td>Single 5</td>
<td>50</td>
</tr>
<tr>
<td>Three 1s</td>
<td>300</td>
</tr>
<tr>
<td>Three of any other number</td>
<td>100 * number</td>
</tr>
<tr>
<td>4 of any number</td>
<td>1000</td>
</tr>
<tr>
<td>5 of any number</td>
<td>2000</td>
</tr>
<tr>
<td>6 of any number</td>
<td>3000</td>
</tr>
<tr>
<td>4 of any number and 2 of another</td>
<td>1500</td>
</tr>
<tr>
<td>1-6 straight</td>
<td>1500</td>
</tr>
<tr>
<td>3 pairs</td>
<td>1500</td>
</tr>
<tr>
<td>2 triplets</td>
<td>2500</td>
</tr>
</tbody>
</table>
<p>The winner is the first person to score 10,000 points. Being the competitive person I am, I decided to give the game some analysis and see if I could determine the expected value for a roll in order to figure out if I should roll or not. I chose to use Go for this because I figured I&rsquo;d be doing a good amount of brute forcing, so I wanted a pretty fast language that I was familiar with. I also just like writing Go.</p>
<h1 id="pure-brute-force">Pure Brute Force</h1>
<p>At first, my idea was to do 2 simple things to generate expected value:</p>
<ol>
<li>Use backtracking recursion to generate all the possible rolls</li>
<li>Find all the ways to partition the dice in each roll and take the maximum scoring partition</li>
</ol>
<p>Starting with my backtracking method, I quickly wrote the following</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">backtrack</span>(<span style="color:#a6e22e">rolls</span> []<span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">toRoll</span> <span style="color:#66d9ef">int</span>) <span style="color:#66d9ef">float64</span> {
	<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">toRoll</span> <span style="color:#f92672">==</span> <span style="color:#ae81ff">0</span> {
		<span style="color:#66d9ef">return</span> float64(<span style="color:#a6e22e">maxScore</span>(<span style="color:#a6e22e">rolls</span>))
	}

	<span style="color:#a6e22e">sum</span> <span style="color:#f92672">:=</span> <span style="color:#ae81ff">0.0</span>
	<span style="color:#66d9ef">for</span> <span style="color:#a6e22e">i</span> <span style="color:#f92672">:=</span> <span style="color:#ae81ff">1</span>; <span style="color:#a6e22e">i</span> <span style="color:#f92672">&lt;=</span> <span style="color:#ae81ff">6</span>; <span style="color:#a6e22e">i</span><span style="color:#f92672">++</span> {
		<span style="color:#a6e22e">sum</span> <span style="color:#f92672">+=</span> <span style="color:#a6e22e">backtrack</span>(append(<span style="color:#a6e22e">rolls</span>, <span style="color:#a6e22e">i</span>), <span style="color:#a6e22e">toRoll</span><span style="color:#f92672">-</span><span style="color:#ae81ff">1</span>)
	}
	<span style="color:#66d9ef">return</span> <span style="color:#a6e22e">sum</span> <span style="color:#f92672">/</span> <span style="color:#ae81ff">6</span>
}
</code></pre></div><p>The idea here is fairly simple: I loop through all 6 possible die rolls, sum them up, and average them. When I&rsquo;ve used up all of my rolls, I score the resulting dice.</p>
<p>Scoring was not quite as simple as I had hoped, though. Partitioning a set is a little complicated, especially in a language like Go with no real concept of set algebra. The general algorithm I found is as follows:</p>
<ol>
<li>If there is only one element, return the set containing just that element. Otherwise, do the following:</li>
<li>Remove the first element from the set and generate the partitions for the remaining set.</li>
<li>For each generated partition, do the following
<ol>
<li>Add the partition containing {removed element} + generated partition</li>
<li>For each set in the generated partition:
<ol>
<li>Add the partition containing {removed element + set} + generated partition \ set</li>
</ol>
</li>
</ol>
</li>
</ol>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">generatePartitions</span>(<span style="color:#a6e22e">rolls</span> []<span style="color:#66d9ef">int</span>) [][][]<span style="color:#66d9ef">int</span> {
	<span style="color:#66d9ef">if</span> len(<span style="color:#a6e22e">rolls</span>) <span style="color:#f92672">==</span> <span style="color:#ae81ff">1</span> {
		<span style="color:#66d9ef">return</span> [][][]<span style="color:#66d9ef">int</span>{{{<span style="color:#a6e22e">rolls</span>[<span style="color:#ae81ff">0</span>]}}}
	}

	<span style="color:#a6e22e">firstElem</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">rolls</span>[<span style="color:#ae81ff">0</span>]
	<span style="color:#a6e22e">rest</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">generatePartitions</span>(<span style="color:#a6e22e">rolls</span>[<span style="color:#ae81ff">1</span>:])
	<span style="color:#a6e22e">toReturn</span> <span style="color:#f92672">:=</span> [][][]<span style="color:#66d9ef">int</span>{}

	<span style="color:#66d9ef">for</span> <span style="color:#a6e22e">_</span>, <span style="color:#a6e22e">elem</span> <span style="color:#f92672">:=</span> <span style="color:#66d9ef">range</span> <span style="color:#a6e22e">rest</span> {
		<span style="color:#a6e22e">toReturn</span> = append(<span style="color:#a6e22e">toReturn</span>, append(<span style="color:#a6e22e">elem</span>, []<span style="color:#66d9ef">int</span>{<span style="color:#a6e22e">firstElem</span>}))

		<span style="color:#66d9ef">for</span> <span style="color:#a6e22e">i</span>, <span style="color:#a6e22e">set</span> <span style="color:#f92672">:=</span> <span style="color:#66d9ef">range</span> <span style="color:#a6e22e">elem</span> {
			<span style="color:#a6e22e">removed</span> <span style="color:#f92672">:=</span> make([][]<span style="color:#66d9ef">int</span>, len(<span style="color:#a6e22e">elem</span>))
			copy(<span style="color:#a6e22e">removed</span>, <span style="color:#a6e22e">elem</span>[:])
			<span style="color:#a6e22e">removed</span> = append(<span style="color:#a6e22e">removed</span>[:<span style="color:#a6e22e">i</span>], <span style="color:#a6e22e">removed</span>[<span style="color:#a6e22e">i</span><span style="color:#f92672">+</span><span style="color:#ae81ff">1</span>:]<span style="color:#f92672">...</span>)
			<span style="color:#a6e22e">toReturn</span> = append(<span style="color:#a6e22e">toReturn</span>, append(<span style="color:#a6e22e">removed</span>, append(<span style="color:#a6e22e">set</span>, <span style="color:#a6e22e">firstElem</span>)))
		}
	}

	<span style="color:#66d9ef">return</span> <span style="color:#a6e22e">toReturn</span>
}
</code></pre></div><p>Then, I take all these partitions, score each one, and return the max of all the scores:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">maxScore</span>(<span style="color:#a6e22e">rolls</span> []<span style="color:#66d9ef">int</span>) <span style="color:#66d9ef">int</span> {
	<span style="color:#a6e22e">max</span> <span style="color:#f92672">:=</span> <span style="color:#ae81ff">0</span>
	<span style="color:#a6e22e">partitions</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">generatePartitions</span>(<span style="color:#a6e22e">rolls</span>)
	<span style="color:#66d9ef">for</span> <span style="color:#a6e22e">_</span>, <span style="color:#a6e22e">partition</span> <span style="color:#f92672">:=</span> <span style="color:#66d9ef">range</span> <span style="color:#a6e22e">partitions</span> {
		<span style="color:#a6e22e">score</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">scoreSuperSet</span>(<span style="color:#a6e22e">partition</span>)
		<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">score</span> &gt; <span style="color:#a6e22e">max</span> {
			<span style="color:#a6e22e">max</span> = <span style="color:#a6e22e">score</span>
		}
	}

	<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">max</span> &gt; <span style="color:#ae81ff">0</span> {
		<span style="color:#66d9ef">return</span> <span style="color:#a6e22e">curScore</span> <span style="color:#f92672">+</span> <span style="color:#a6e22e">max</span>
	}
	<span style="color:#66d9ef">return</span> <span style="color:#ae81ff">0</span>
}
</code></pre></div><p>After writing this code, however, I realized that it had quite a few problems:</p>
<ul>
<li>Backtracking could have to check up to 6^6 (46,656) rolls</li>
<li>Each roll would have to be partitioned in every way possible. This follows the Bell numbers, so in the worst case, there are 203 possible partitions to check. Combined with the previous number, in the worst case, I have to check 947,1168 partitions.</li>
<li>I&rsquo;m only accounting for one roll! The part of Farkle that makes it so hard to analyze is the ability to roll multiple times, but if you take a look at my code, I&rsquo;m only checking the possibilities for the direct next roll.</li>
</ul>
<p>Let&rsquo;s make a checklist for things to fix:</p>
<ul>
<li><input disabled="" type="checkbox"> Be smarter than backtracking all possibilities</li>
<li><input disabled="" type="checkbox"> Find a way to score a roll in constant time</li>
<li><input disabled="" type="checkbox"> Account for possibly making multiple more rolls</li>
</ul>
<h1 id="improving-scoring">Improving Scoring</h1>
<p>Something didn&rsquo;t sit right with me regarding how I did scoring. It seemed totally wrong that I would need to generate up to 203 combinations and take their max score. After all, I don&rsquo;t consider all possible combinations when scoring my rolls playing with my family. I group the dice according to their numbers and see if the groups fit into any of the scoring categories. Turns out, that works pretty well for real scoring. Here&rsquo;s what I came up with:</p>
<p>1. Go through all the dice and record how many of each number there are. Also record which dice appear with which frequency</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">scoreRoll</span>(<span style="color:#a6e22e">rolls</span> []<span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">score</span> <span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">depth</span> <span style="color:#66d9ef">int</span>) <span style="color:#66d9ef">float64</span> {
  <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">occurences</span> [<span style="color:#ae81ff">6</span>]<span style="color:#66d9ef">int</span>
  <span style="color:#66d9ef">for</span> <span style="color:#a6e22e">_</span>, <span style="color:#a6e22e">elem</span> <span style="color:#f92672">:=</span> <span style="color:#66d9ef">range</span> <span style="color:#a6e22e">rolls</span> {
    <span style="color:#a6e22e">occurences</span>[<span style="color:#a6e22e">elem</span><span style="color:#f92672">-</span><span style="color:#ae81ff">1</span>]<span style="color:#f92672">++</span>
  }
  <span style="color:#66d9ef">var</span> <span style="color:#a6e22e">frequencies</span> [<span style="color:#ae81ff">7</span>][]<span style="color:#66d9ef">int</span>

  <span style="color:#66d9ef">for</span> <span style="color:#a6e22e">i</span> <span style="color:#f92672">:=</span> <span style="color:#ae81ff">0</span>; <span style="color:#a6e22e">i</span> &lt; <span style="color:#ae81ff">6</span>; <span style="color:#a6e22e">i</span><span style="color:#f92672">++</span> {
    <span style="color:#a6e22e">occurs</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#a6e22e">i</span>]
    <span style="color:#a6e22e">frequencies</span>[<span style="color:#a6e22e">occurs</span>] = append(<span style="color:#a6e22e">frequencies</span>[<span style="color:#a6e22e">occurs</span>], <span style="color:#a6e22e">i</span>)
  }
<span style="color:#f92672">...</span>
</code></pre></div><p>2. Compare the frequencies against the rules of the game</p>
<p>I won&rsquo;t include this part since it&rsquo;s just a series of if/else statements. You can find the code at this iteration <a href="https://github.com/Pwpon500/farkle-odds/blob/fb544b026a537c3f3d75a7d9c08e871c545c0b04/main.go">here</a> if interested.</p>
<p>3. Count the number of remaining 1s and 5s</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">if</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#ae81ff">0</span>] &lt; <span style="color:#ae81ff">3</span> <span style="color:#f92672">&amp;&amp;</span> <span style="color:#a6e22e">countIndiv</span> {
<span style="color:#a6e22e">toReturn</span> <span style="color:#f92672">+=</span> float64(<span style="color:#ae81ff">100</span> <span style="color:#f92672">*</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#ae81ff">0</span>])
<span style="color:#a6e22e">numUsed</span> <span style="color:#f92672">+=</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#ae81ff">0</span>]
}
<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#ae81ff">4</span>] &lt; <span style="color:#ae81ff">3</span> <span style="color:#f92672">&amp;&amp;</span> <span style="color:#a6e22e">countIndiv</span> {
<span style="color:#a6e22e">toReturn</span> <span style="color:#f92672">+=</span> float64(<span style="color:#ae81ff">50</span> <span style="color:#f92672">*</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#ae81ff">4</span>])
<span style="color:#a6e22e">numUsed</span> <span style="color:#f92672">+=</span> <span style="color:#a6e22e">occurences</span>[<span style="color:#ae81ff">4</span>]
}
</code></pre></div><p>4. If there is no score, the player gets no points. Otherwise, they get their scored points plus the score they already have</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">if</span> <span style="color:#a6e22e">toReturn</span> <span style="color:#f92672">==</span> <span style="color:#ae81ff">0</span> {
	<span style="color:#66d9ef">return</span> <span style="color:#ae81ff">0</span>
}
<span style="color:#66d9ef">return</span> <span style="color:#a6e22e">toReturn</span> <span style="color:#f92672">+</span> float64(<span style="color:#a6e22e">score</span>)
</code></pre></div><p>Now, I have a scoring algorithm that will work in constant time. Let&rsquo;s update that checklist:</p>
<ul>
<li><input disabled="" type="checkbox"> Be smarter than backtracking all possibilities</li>
<li><input checked="" disabled="" type="checkbox"> Find a way to score a roll in constant time</li>
<li><input disabled="" type="checkbox"> Account for possibly making multiple more rolls</li>
</ul>
<h1 id="more-backtracking">More Backtracking</h1>
<p>Taking extra rolls into account is actually pretty simple in theory. When you score each roll, take the expected value of rolling again. If that expected value is higher than what you&rsquo;ve already scored, use it. Otherwise, keep your current score. This can be simply implemented as follows:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">expectedRoll</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">backtrack</span>([]<span style="color:#66d9ef">int</span>{}, <span style="color:#a6e22e">numLeft</span>, int(<span style="color:#a6e22e">toReturn</span>))
<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">expectedRoll</span> &gt; float64(<span style="color:#a6e22e">toReturn</span>) {
	<span style="color:#a6e22e">toReturn</span> = <span style="color:#a6e22e">expectedRoll</span>
}
</code></pre></div><p>I ran this code and waited. I waited some more, and I waited even longer. Looking back, I committed the cardinal sin of recursion: recursion without a base case. Each roll would check the expected value of another roll, and the cycle would infinitely repeat. To fix this, I added a simple <code>depth</code> parameter to backtrack and implemented it as follows:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">backtrack</span>(<span style="color:#a6e22e">rolls</span> []<span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">toRoll</span> <span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">score</span> <span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">depth</span> <span style="color:#66d9ef">int</span>) <span style="color:#66d9ef">float64</span> {
	<span style="color:#66d9ef">if</span> <span style="color:#a6e22e">depth</span> <span style="color:#f92672">&gt;=</span> <span style="color:#a6e22e">maxDepth</span> {
		<span style="color:#66d9ef">return</span> <span style="color:#ae81ff">0</span>
	}
</code></pre></div><p><code>maxDepth</code> is just a global variable that the user sets to define how deep they want to recurse. I then pass depth into my scoring function and change my expected roll call to implement depth:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#a6e22e">expectedRoll</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">backtrack</span>([]<span style="color:#66d9ef">int</span>{}, <span style="color:#a6e22e">numLeft</span>, int(<span style="color:#a6e22e">toReturn</span>), <span style="color:#a6e22e">depth</span><span style="color:#f92672">+</span><span style="color:#ae81ff">1</span>)
</code></pre></div><p>Now, theoretically, I have an algorithm that&rsquo;ll go to whatever arbitrary depth I want. However, upon running this, I can only go to a depth of 2 before the program explodes running time. Since I&rsquo;m backtracking, each added level of depth makes the running time go up exponentially. So I&rsquo;ve fixed one problem, but just amplified another. Let&rsquo;s update the checklist:</p>
<ul>
<li><input disabled="" type="checkbox"> Be smarter than backtracking all possibilities</li>
<li><input checked="" disabled="" type="checkbox"> Find a way to score a roll in constant time</li>
<li><input checked="" disabled="" type="checkbox"> Account for possibly making multiple more rolls</li>
</ul>
<h1 id="finding-a-model">Finding A Model</h1>
<p>Faced with no idea for how to do this without backtracking, I had the idea of dynamic programming. If I could just find a decent model for the expected values, I could get rid of all the costly backtracking. I added support for command line flags and generated all the expected values from 1 to 400 with the following command (this is in fish but a similar thing can be done in bash):</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-bash" data-lang="bash"><span style="color:#66d9ef">for</span> x in <span style="color:#f92672">(</span>seq 400<span style="color:#f92672">)</span>
    ./farkle-odds -dice <span style="color:#ae81ff">1</span> -score $x
end
</code></pre></div><p>I used plot.ly to plot all this out and saw something pretty amazing:</p>
<p><img src="/img/depth1_farkle.png" alt="Farkle at depth 1">{:width=&quot;750px&rdquo;}</p>
<p>It&rsquo;s linear! Realizing this, I generated lines of best fit for each of the numbers of dice. I then added the <code>approxScore</code> method in to use the lines of best fit to approximate the expected value of a score:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">approxScore</span>(<span style="color:#a6e22e">toRoll</span> <span style="color:#66d9ef">int</span>, <span style="color:#a6e22e">score</span> <span style="color:#66d9ef">int</span>) <span style="color:#66d9ef">float64</span> {
	<span style="color:#66d9ef">return</span> <span style="color:#a6e22e">mVals</span>[<span style="color:#a6e22e">toRoll</span><span style="color:#f92672">-</span><span style="color:#ae81ff">1</span>]<span style="color:#f92672">*</span>float64(<span style="color:#a6e22e">score</span>) <span style="color:#f92672">+</span> <span style="color:#a6e22e">bVals</span>[<span style="color:#a6e22e">toRoll</span><span style="color:#f92672">-</span><span style="color:#ae81ff">1</span>]
}
</code></pre></div><h1 id="generating-lines">Generating Lines</h1>
<p>After some plotting of greater depths, I concluded that the greater depths are <em>almost</em> linear, and any deviations from a linear model were not significant enough to substantially impact the generated expected values. Using this linearity to my advantage, I used the following algorithm to generate coefficients for a linear model for any arbitrary depth:</p>
<ol>
<li>Use <code>backtrack</code> to find 2 expected values</li>
<li>Use simple algebra to find the <code>m</code> and <code>b</code> values for the line between those points</li>
<li>Record those <code>m</code> and <code>b</code> values in JSON to a file</li>
<li>Replace the old <code>m</code> and <code>b</code> values with the new ones and repeat the process</li>
</ol>
<p>Here&rsquo;s what that looks like in actual code:</p>
<div class="highlight"><pre style="color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4"><code class="language-go" data-lang="go"><span style="color:#66d9ef">func</span> <span style="color:#a6e22e">generateCoeffs</span>() {
	<span style="color:#66d9ef">for</span> <span style="color:#a6e22e">i</span> <span style="color:#f92672">:=</span> <span style="color:#ae81ff">1</span>; <span style="color:#a6e22e">i</span> <span style="color:#f92672">&lt;=</span> <span style="color:#ae81ff">50</span>; <span style="color:#a6e22e">i</span><span style="color:#f92672">++</span> {
		<span style="color:#a6e22e">newM</span> <span style="color:#f92672">:=</span> []<span style="color:#66d9ef">float64</span>{}
		<span style="color:#a6e22e">newB</span> <span style="color:#f92672">:=</span> []<span style="color:#66d9ef">float64</span>{}
		<span style="color:#66d9ef">for</span> <span style="color:#a6e22e">j</span> <span style="color:#f92672">:=</span> <span style="color:#ae81ff">1</span>; <span style="color:#a6e22e">j</span> <span style="color:#f92672">&lt;=</span> <span style="color:#ae81ff">6</span>; <span style="color:#a6e22e">j</span><span style="color:#f92672">++</span> {
			<span style="color:#75715e">// find high and low vals for this iteration
</span><span style="color:#75715e"></span>			<span style="color:#a6e22e">lowVal</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">backtrack</span>([]<span style="color:#66d9ef">int</span>{}, <span style="color:#a6e22e">j</span>, <span style="color:#ae81ff">0</span>)
			<span style="color:#a6e22e">highVal</span> <span style="color:#f92672">:=</span> <span style="color:#a6e22e">backtrack</span>([]<span style="color:#66d9ef">int</span>{}, <span style="color:#a6e22e">j</span>, <span style="color:#ae81ff">400</span>)
			<span style="color:#75715e">// use high and low vals to generate line
</span><span style="color:#75715e"></span>			<span style="color:#a6e22e">newM</span> = append(<span style="color:#a6e22e">newM</span>, (<span style="color:#a6e22e">highVal</span><span style="color:#f92672">-</span><span style="color:#a6e22e">lowVal</span>)<span style="color:#f92672">/</span><span style="color:#ae81ff">400</span>)
			<span style="color:#a6e22e">newB</span> = append(<span style="color:#a6e22e">newB</span>, <span style="color:#a6e22e">lowVal</span>)
		}
		<span style="color:#a6e22e">coeffs</span>.<span style="color:#a6e22e">MCoeffs</span> = <span style="color:#a6e22e">newM</span>
		<span style="color:#a6e22e">coeffs</span>.<span style="color:#a6e22e">BCoeffs</span> = <span style="color:#a6e22e">newB</span>
		<span style="color:#a6e22e">writeVals</span>(<span style="color:#a6e22e">newM</span>, <span style="color:#a6e22e">newB</span>, <span style="color:#a6e22e">i</span>)
	}
	<span style="color:#a6e22e">fmt</span>.<span style="color:#a6e22e">Println</span>(<span style="color:#e6db74">&#34;Coefficients generated and written&#34;</span>)
}
</code></pre></div><p>Now, we can finally check off that last box!</p>
<ul>
<li><input checked="" disabled="" type="checkbox"> Be smarter than backtracking all possibilities</li>
<li><input checked="" disabled="" type="checkbox"> Find a way to score a roll in constant time</li>
<li><input checked="" disabled="" type="checkbox"> Account for possibly making multiple more rolls</li>
</ul>
<h1 id="conclusion">Conclusion</h1>
<p>After some refactoring, I&rsquo;m really happy with the result of this little experiment. The linear model isn&rsquo;t a perfect fit for the expected values, but it&rsquo;s damn good and takes constant time to run.</p>
<p>If you&rsquo;re interested in seeing what the code looks like, it can be found <a href="https://github.com/Pwpon500/farkle-odds">here</a>. Make sure to let me know if there are any improvements I can make.</p>
<p>Also, here is some of the plotted data of the expected values for different dice rolls:</p>
<p><a href="https://plot.ly/~Pwpon500/1">Depth 1</a> <a href="https://plot.ly/~Pwpon500/3">Depth 2</a></p>
]]></content></item><item><title>VPLS with OpenBSD</title><link>https://pawa.lt/posts/2018/01/vpls-with-openbsd/</link><pubDate>Mon, 15 Jan 2018 00:00:00 +0000</pubDate><guid>https://pawa.lt/posts/2018/01/vpls-with-openbsd/</guid><description>VPLS is extremely useful in allowing multiple sites to be connected to a single bridged domain. Unfortunately, VPLS networks are typically implemented with proprietary technology like a Cisco router. OpenBSD lets us break free of the typical restrictions of proprietary technology and use 100% free software to make a full-fledged VPLS network.
Why VPLS? With VPLS, you can deliver a layer 2 circuit over a routed backbone. This lets you extend that circuit from one location to another over a route that you can easily control.</description><content type="html"><![CDATA[<p>VPLS is extremely useful in allowing multiple sites to be connected to a single bridged domain. Unfortunately, VPLS networks are typically implemented with proprietary technology like a Cisco router. OpenBSD lets us break free of the typical restrictions of proprietary technology and use 100% free software to make a full-fledged VPLS network.</p>
<h2 id="why-vpls">Why VPLS?</h2>
<p>With VPLS, you can deliver a layer 2 circuit over a routed backbone. This lets you extend that circuit from one location to another over a route that you can easily control. If the backbone is not routed, you leave the path-finding up to spanning tree; while that works for smaller networks, the bigger your layer 2 domain gets, the less-reliable spanning tree gets. Eventually, it may start taking spaghetti-like paths, and there will be little you can do about it.</p>
<p>VPLS is also useful in allowing the extension of layer 2 domain to a remote site. For example, in the case of VoIP phones, VPLS would allow the phones to be directly connected to the remote VoIP server, requiring no configuration whatsoever at the client site.</p>
<h2 id="vpls-on-a-high-level">VPLS on a High Level</h2>
<p>VPLS works by creating MPLS pseudowires (point-to-point layer 2 circuits) to every node you want to be part of your bridged domain. You then add each pseudowire as well as the physical interface you want to be &ldquo;bridged into&rdquo; the VPLS domain. In basic terms, you have now created a &ldquo;virtual switch,&rdquo; plugged each pseudowire into it, and plugged your physical interface into it.</p>
<h2 id="the-setup">The Setup</h2>
<p>This setup is completely virtualized, but it can be replicated easily with physical servers as well. There will be one provider and three provider edges. Each router will have a unique router-id by which it will be identified in OSPF and LDP. That address will be assigned to a secondary loopback interface (lo1) and advertised to the other routers using OSPF. The setup looks like this:
<img src="/img/vpls_openbsd_1.png" alt="OpenBSD VPLS Base Setup">{:width=&quot;500px&rdquo;}</p>
<p>Each PE will be directly attached to the provider with a /30. I&rsquo;m using the 172.30.2.X IP scheme, but you can use whatever you want. Just make sure to match up those addresses between the provider and PE. Similarly, my use of 10.0.0.X for the router-id&rsquo;s can changed to whatever you want.</p>
<h2 id="global-configuration">Global Configuration</h2>
<p>On each of the nodes, you will have to enable some services. I&rsquo;m doing this in /etc/rc.conf.local. Append these two lines at the bottom to enable OSPF and LDP. Also make sure to wipe any lines disabling OSPF or LDP out of /etc/rc.conf and /etc/rc.conf.local:</p>
<pre><code>ospfd_flags=&quot;&quot;
ldpd_flags=&quot;&quot;
</code></pre><p>You&rsquo;ll also need to give each router its appropriate router-id on its loopback interface. You can put this in /etc/rc.local or use the hostname.XXX format. I prefer the hostname.XXX format.</p>
<p>/etc/hostname.lo1:</p>
<pre><code>inet 10.0.0.1 255.255.255.255
description &quot;id_loopback&quot;
</code></pre><p>Make sure to change that 10.0.0.X address for each router.</p>
<h2 id="provider-configuration">Provider Configuration</h2>
<p>The provider is the simplest node to configure since it just acts as a glorified label switch.</p>
<p>/etc/ospfd.conf:</p>
<pre><code>router-id 10.0.0.1

area 0.0.0.0 {
    interface lo1
    interface re0
    interface re1
    interface re2
}
</code></pre><p>/etc/ldpd.conf:</p>
<pre><code>router-id 10.0.0.1

address-family ipv4 {
    interface re0
    interface re1
    interface re2
}
</code></pre><p>/etc/hostname.re0:</p>
<pre><code>inet 172.30.2.1 255.255.255.252
mpls
description &quot;p1_edge&quot;
</code></pre><p>It&rsquo;s critical that you have the <code>mpls</code> line in your configuration. This lets OpenBSD know to treat that interface as a provider-facing interface. Repeat this interface configuration for each PE. Make sure to use a different IP for each interface as well as a different subnet. I recommend using consecutive /30 blocks (172.30.2.0/30, 172.30.2.4/30, 172.30.2.8/30).</p>
<h2 id="pe-configuration">PE Configuration</h2>
<p>The configuration of PEs is similar to that of the provider, but pseudowires have to be created to each other PE. Each PE has two physical interfaces. re0 is the provider-facing interface, and re1 is the client-facing interface.</p>
<p>/etc/ospfd.conf:</p>
<pre><code>router-id 10.0.0.2

area 0.0.0.0 {
    interface lo1
    interface re0
}
</code></pre><p>First, we have to create our pseudowires. These follow the mpwX naming convention. All we have to do in /etc/hostname.XXX is create them and bring them up. LDPD takes care of the rest. We also have to bring up our physical client-facing interface.</p>
<p>/etc/hostname.mpw0:</p>
<pre><code>create
up
</code></pre><p>/etc/hostname.re1:</p>
<pre><code>up
</code></pre><p>We also need to bring up our bridge interface and add our pseudowires and our physical interface to it.</p>
<p>/etc/hostname.bridge0:</p>
<pre><code>add re1
add mpw0
add mpw1
up
description &quot;vpls_bridge&quot;
</code></pre><p>Remember that in this LDP configuration, re0 is provider-facing, and re1 is client-facing. If you want some extra information on how ldpd.conf works, check out the <a href="https://man.openbsd.org/ldpd.conf.5">OpenBSD ldpd.conf man page</a>.
/etc/ldpd.conf:</p>
<pre><code>router-id 10.0.0.2

address-family ipv4 {
    interface re0
}

l2vpn pe1 type vpls {
    bridge bridge0
    interface re1

    pseudowire mpw0 {
        neighbor-id 10.0.0.3
        pw-id 100
    }
    pseudowire mpw1 {
        neighbor-id 10.0.0.4
        pw-id 100
    }
}
</code></pre><p>Repeat this process with the rest of the PEs, changing the appropriate IPs and router-id&rsquo;s. After that&rsquo;s all done, restart ldpd and ospfd, use /etc/netstart to bring up interfaces, and you&rsquo;re ready to go!</p>
<p>You can find my full configuration <a href="https://github.com/Pwpon500/vpls-openbsd">here</a>. Use this if you need any extra help figuring out the configurations for the nodes.</p>
<h2 id="testing">Testing</h2>
<p>To test your VPLS setup, connect clients to your physical interfaces on all your PEs, and give them all static IPs in the same subnet. Try to ping each other and see if the pings return. If they do, you did it! You&rsquo;ve now created a functional VPLS network.</p>
<h2 id="diagnostics">Diagnostics</h2>
<p>More likely, however, your network doesn&rsquo;t work. Don&rsquo;t worry! OpenBSD has some great tools for viewing LDP status. Here are some common commands and what they look like on a working PE:</p>
<pre><code>pe1 / root / 23:18:29
&gt; ~ # ldpctl show neighbor
AF   ID              State       Remote Address    Uptime
ipv4 10.0.0.1        OPERATIONAL 10.0.0.1        00:03:53
ipv4 10.0.0.3        OPERATIONAL 10.0.0.3        00:03:11
ipv4 10.0.0.5        OPERATIONAL 10.0.0.5        00:03:11
pe1 / root / 23:18:29
&gt; ~ # ldpctl show l2vpn pseudowire
Interface   Neighbor        PWID           Status
mpw0        10.0.0.3        100            UP
mpw1        10.0.0.4        100            UP
</code></pre><p>You can also use ospfctl to if you think the problem may lie in the routing.</p>
<h2 id="conclusion">Conclusion</h2>
<p>VPLS is some amazing technology, and hopefully, you can implement it yourself with the help of this post.</p>
]]></content></item></channel></rss>