Hello World
Introducing my research blog.
Hello world!
I study theoretical and applied machine learning - learning efficiency, concept representation, and information compression. My work sits at the intersection of the statistical learning theory and machine learning systems.
My intuition is that theoretical models can extend to empirical results and informative answers come from abstracting experimental conclusions into concrete theorems and proofs. The questions I find compelling are ones that reach across both research areas: how do models develop internal connections over the course of training? What properties of data, architecture, and optimization determine what is learned? How do the representations that emperical research uncover relate to the scaling behavior that learning theory models? And how do these abstractions hold up when confronted with the constraints of real systems - memory, throughput, latency, and scale?
One of the many joys of machine learning research is the compatability of theory and experimentation. Theoretical ideas are can be cheaply tested (relatively speaking) on widely accessible hardware without bureaucratic processes. This solves an important problem: too many theoretical ideas yield zero practical applications, and many experiments produce unexplainable results. The ability to quantify tangible outcomes and connect these observations to mathematical foundations provides computer scientists with a unique ability: to simplify the complexity of universe's boundless information.
This blog is meant to be a working notebook. Some posts will be sketches of arguments I am still trying to formalize. Others will be experimental notes, plots, conjectures, and anomalies that I have not yet been able to fully explain. Occasionally there will be more developed pieces that bring theory and application together under one umbrella. None of these are intended as substitutes for full papers and candidly, I expect some of my conjectures to be incorrect and revised. The goal is to make my entire research process visible.
These notes are written for other researchers interested in machine learning, though I hope they remain accessible to anyone curious about theoretical learning and why machine learning systems behave the way they do. If something here is useful, wrong, or incomplete, I would be glad to hear about it. Progress on these questions is not an individual undertaking; we all stand on the shoulders of giants.
Thanks for reading.