Extracting a web page’s title with command-line tools and the pup HTML parser instead of relying on an external title service.
Making the Hugo → S3 upload process much more efficient by tracking file hashes.
Dealing with PUNCT nodes in interlinear glossing.
A CSS-based experiment in displaying words, lemmas, and parts of speech as three-line interlinear text for a language-learning project.
leipzig.js is a library for applying
interlinear gloss to texts for linguistic analysis. In this post, I experiment a little with this libary to evaluate whether it would work for a little project of mine.
Starting a new devlog about Hedghog, a new language learning app and some thoughts about the interlinear display of lemmas.
Splitting text into sentences is one of those tasks that looks simple but on closer inspection is more difficult than you think. A common approach is to use regular expressions to divide up the text on punction marks. But without adding layers of complexity, that method fails on some sentences. This is a method using
spaCy.
Why variables changed inside a Bash pipeline may remain unchanged outside it, and how the lastpipe option alters that behaviour.
A Hazel workflow that extracts text from financial-statement PDFs, renames them consistently, and imports them into DEVONthink with tags.
Recursively renaming DEVONthink tags with AppleScript to simplify a hierarchical tagging system and change its separators.