It started earlier in the year, but June is when I actually got into it. July is when I decided to commit. So the clock runs 1 July 2026 to 1 July 2031.
I picked five years because it is long enough to get good at something hard and short enough that I cannot keep putting it off. The number in the corner is days gone and days left. I did not label it. I wanted to see it every day without explaining myself to anyone, including me.
Here is what it is for.
The north star
Designing systems that remember, reason about, and act on real-world entities under uncertainty.
That sentence is doing a lot of work, so let me unpack it.
Most software rests on a quiet assumption about identity. People become rows in tables. Places become strings or coordinates. Companies become UUIDs. While humans stay in the loop, those shortcuts mostly hold. Someone notices a record looks wrong and fixes it.
As decisions shift to automation, they stop holding. Some of those failures go unnoticed for a long time. All of them get expensive.
So the question I keep circling is how identity, relationships, context and evidence come together well enough that a system can safely answer four basic things. What is this? Have we seen it before? Can we act on it? Is it safe to?
Where I am coming from
I do not come at this from research. I come at it from the side where being wrong costs something.
Most of my work sits close to the business, where a number has to be right and consistent, and where someone is about to make a decision on it. Attribution models. Customer records. Fragmented identifiers. You spend enough time explaining why a figure moved, and you find that it usually moved because the data was quietly wrong about who or what it was describing.
That is a useful place to stand. It is much harder to hand-wave a bad match when a person is waiting on the answer.
The pieces I am working through
Not a syllabus, just the things that keep turning out to matter.
Entity resolution. Deciding whether two records are the same thing. It is old, well understood in theory, and still broken in practice everywhere I look, mostly because people treat it as a matching problem rather than a representation problem.
Knowledge graphs. Less the technology, more the discipline of writing down what relates to what, and being honest about how you know.
Latent spaces. Embeddings are the dominant way machines represent meaning right now. I want to understand what they are actually good at and, more importantly, where they quietly lose the thing you cared about.
Reinforcement learning. Because acting under uncertainty is a different problem from predicting under uncertainty, and the second one is where all the attention has gone.
Fingerprints and provenance. Knowing that a thing is what it was, and being able to show it. This is the least glamorous item on the list and probably the most useful. Most arguments about trust in systems turn out, on inspection, to be arguments about whether anyone can tell what changed and who changed it. Every page on this site carries a fingerprint for exactly this reason, and I have written up what it does and does not prove.
Man and machine, working together. Licklider wrote about human-computer symbiosis in 1960 and the framing still holds up better than most of what has been written since. Not the machine replacing the judgment. The machine handling what it is genuinely better at, and knowing where its part ends.
Whose data is it
There is a version of this that is not technical at all.
Right now, using almost any system means handing over a copy of yourself and hoping. You do not lend your data. You surrender it, permanently, to a place you cannot see, and the relationship is over the moment you close the tab, except your data stays behind.
Tim Berners-Lee has spent years arguing it could work the other way round. You keep your own data. Systems ask for the parts they need, for as long as the relationship lasts, and you can end it. Decentralised identity and self-sovereign identity are the names people give to different shapes of that idea.
I find the picture genuinely appealing. Imagine choosing to give a platform more of yourself precisely because you intend to stay a while, and being able to take it back when you do not.
I am also aware of how hard it is. Somebody has to host the thing. Most people will not run their own. And the moment your data is portable, the question of whether a system has correctly identified you stops being an internal detail and becomes the whole basis of the exchange.
Which is exactly the problem I already work on, arriving from a different direction.
Signatures, and why they have to last
If a decision about identity is going to be defensible later, it has to be verifiable later, by someone who does not trust you.
That is a cryptography question, and it is the one part of cryptography I want to be genuinely competent in. I am not trying to become a cryptographer. I care about three properties and mostly stop there:
- Non-repudiation. The party who made the claim cannot later deny making it.
- Confidentiality. Only the intended party can read it.
- Security. It holds up against someone actively trying to break it, not just against noise.
I recently joined the PKI Consortium as a member, largely to learn this properly from people who do it for a living, and hopefully to contribute something back on where it meets representation and identity.
There is a deadline hiding in here. A signature made today may need to verify in ten or fifteen years, and the algorithms most signatures currently use are the ones a sufficiently capable quantum computer would break. The usual framing is “harvest now, decrypt later”, which is about confidentiality. The version that worries me more is quieter: every long-lived signed record we are creating today is being created with a shelf life nobody has priced in.
I do not think I will contribute to post-quantum cryptography itself. I would like to understand it well enough to know which of the things I build will still be standing.
Which brings me to the thing everyone wants to talk about.
On singularity, and why I am not writing about it
The singularity argument is that once a system can improve itself, improvement compounds and you get a step change that nobody is positioned to steer. It is a serious idea. Good and careful people take it seriously.
My honest position is that it is above my pay grade and outside my control, and that arguing about it is more fun than useful.
What is inside my control is much smaller. Whether the system in front of me knows what it is acting on. Whether it knows how sure it is. Whether it stops when it should.
If the big version ever arrives, those questions do not get less important. They get more important, and somebody will wish they had been answered earlier.
What I actually want to build
Agents should act only on what they understand, remember, and trust.
If I am right, the systems worth having in a few years will:
- know what entity they are acting on
- know why they believe that
- know how confident they are
- refuse to act when confidence is not enough
- leave a trail a human can inspect
That is not AGI. I am not chasing that.
To me this mostly feels like common sense. Responsible autonomy is a tidy name for it, and I am slightly suspicious of tidy names, because they have a way of becoming marketing before they become real. But it is a smaller and more useful target than the big version, especially alongside a day job and a life.
How I will know
Five years is long enough to fool yourself, so here is what I would count as moving in the right direction:
- a body of work on representation and information quality that other people build on
- a book that was worth writing, and papers on the subject after it
- people I taught or backed who went further than I did, or a community I was genuinely useful in
- one algorithm or open source project I designed, in the open, that someone else found useful
- enough maths to read a paper and implement it without flinching
- contributing to the open conversation, the tools, and speaking about this properly
- leading a team or a project built around it
Some of that is on List 99. Most of it is not, because the list is about a life and this is about a direction.
Open risks
I would rather write these down than pretend they are not there.
Life gets busy. Work, family, everything else. The most likely outcome is not failure, it is that this takes longer than I expect.
Motivation. Doing a day job at a high standard and building something like this at the same time is a real cost. I think I want the challenge. Ask me in year three.
Someone solves it first. A frontier lab may well get there before I do. Genuinely possible. It does not change my resolve much, because I would still want to understand it, pass it on, and honestly, because it is fun.
Outsourcing the learning. This is the one I think about most. Understanding something properly is slow, and there is now a very tempting shortcut that produces the feeling of understanding without the thing itself. Resisting that is hard. Interestingly, doing it well probably means designing a context layer that helps rather than replaces, which is squarely inside what these five years are supposed to be about.
How to help
You already did, by reading this.
I still carry a decent amount of impostor syndrome about all of it. But you can find me on LinkedIn, or just follow the work. I will be doing this alongside some genuinely brilliant people, and I am looking forward to it.