We rewrote the core of BlockBen’s blockchain and balance management system. 30–40× faster, sub-millisecond operations, and an architecture with no single point whose failure could bring the system to a halt. This article explains what happened behind the longer-than-planned system upgrade on August 31, 2026 — and why we had to develop our own protocol to make it happen.

Recently the system upgrade took longer than planned. There is a very good reason for that, and it is not what you might think at first: the code rollout went smoothly, but the AWS infrastructure changes took longer to complete in production than they did in testing. It was probably cloud provider overload. We cannot promise that this will never happen again. What we can promise is that the system we put into production during that process was worth the wait.

If you opened the BlockBen app this morning, you may have noticed that it seemed to load a little faster. You are not imagining it. One of the biggest technological developments of the past few months has now been integrated into the system: we have completely rewritten and rethought the most important components of our blockchain and balance management infrastructure.

This is the kind of development that is difficult to notice from the outside. There is no new button, no new screen, no new animation. Everything is simply faster, more stable, and prepared for significantly higher loads behind the scenes. Now I’ll try to explain what we actually built — in a way that does not require you to be a software engineer, while still showing the engineering work behind it, rather than just throwing a marketing number at you.

What Was Wrong With the Old System?

The BlockBen system has several different “brains”, but I want to focus on two of them here. One is the Ledger: it organizes transactions into blocks, builds the blockchain, and generates the cryptographic proofs that anyone can verify using QantrumScan. I’ll write more about this in another article. The other is the balance management layer, internally called Balancer. It keeps track of account balances, determines whether there is sufficient balance for every transaction, and records the resulting changes.

In the old system, these two were two separate worlds. Separate code, a separate internal language, separate data structures. Both worked, and both could scale — but they were not fast enough, and more importantly in the long run, they had started to diverge. We were solving the same problem in two different places, in two different ways, and had to maintain, test, and debug both implementations separately.

When we sat down to redesign the system, the first realization was not about speed. It was that the two systems worked in essentially the same way internally. Both handled a large dataset that we divided into multiple parts so that multiple machines could work on it in parallel. In both systems, every part was served by multiple machines at the same time, so if one failed, the others could continue. In both cases, these machines elected a leader, and if the leader did not perform properly, it was replaced. And in both systems, a transaction could affect multiple parts at the same time, meaning that either all of them had to execute together, or none of them could.

Two seemingly completely different business functions, following the same internal pattern. That was when we decided not to simply make two systems faster, but to build one common foundation that both could run on. To make this possible, we developed a completely custom protocol. Its internal name: the Phantom protocol. You will hear more about it in the future.

Why Isn’t a Balance Just a Column in a Table?

Many people think keeping track of a balance is simple: there is a number, you add something to it, or subtract something from it. Anyone who has worked with banking systems knows that it is never that simple.

A transaction has a booking date and a value date, and the two are not always the same. The system has to be able to correctly reconstruct both views for any point in the past. There are locked balances and available balances. There are accounts that cannot go negative — customer accounts, for example. And there are accounts where we prohibit the balance from going negative even during a transaction, even if the final result would be positive.

There is also a special type of account that is a good example of where the speed of a financial system is determined. Imagine a revenue account that receives something from practically every transaction. If we treated this account in the same way as a customer account — stopping at every transaction, checking the balance, locking it — then every transaction in the system would have to pass through a single gate, waiting in line one after another. It would be like going into a supermarket with ten checkout counters, only to discover that every customer has to sign a piece of paper with the same single cashier before they can pay.

That is why we handle these accounts differently: transactions can be booked to them in parallel, and the balance is settled periodically rather than after every individual entry. Withdrawals can only be made up to the amount that has already been settled — but that is an accounting question, not a technological one.

The key point is this: Balancer is an accounting engine whose every decision has to be made within milliseconds for each transaction. For this reason, balance checks run entirely in memory rather than through database queries. The “check how much there is, then write” approach is not safe under parallel load — two transactions could see the same balance and both assume that sufficient funds are available. We cannot allow that in a financial system, not even for a millisecond.

Phantom: The Coordinator Doesn’t Need to Remember

Let’s look at the problem more closely. If a transaction affects multiple accounts — for example, a customer account on one unit and a revenue account on another — the two units have to make a joint decision: either both commit, or neither does. For fifty years, computer science has used essentially the same pattern for this: a coordinator first asks everyone whether they are ready, and then announces the decision.

The weak point of this pattern has remained the same for fifty years: the coordinator has to remember. It has to keep a record of whom it asked. If the coordinator crashes between the question and the decision, the participants are left waiting, with locked balances, for a decision that never arrives — unless someone can recover the dead coordinator’s log. The world’s largest systems have built enormous and complicated layers of protection around this problem. They all work. They are all complicated. And they all add latency.

The Reverse Question

We asked the question differently. We did not ask how to protect the coordinator. We asked: why does it need to remember anything at all?

The participants already know whether they were ready. The only missing information is who was asked — and that information can be given to the participants themselves. Imagine a meeting where the secretary does not keep minutes. Instead, every participant gets a small note showing who is sitting at the table. If the secretary disappears, any new secretary can walk in, ask anyone for their note, look around, and immediately know who is missing. There is nothing to search for. Nothing to guess.

In the Phantom protocol, this “note” is just a few bytes that every participant receives. From that point on, there is nothing for the coordinator to lose. It calculates the decision, sends it, and forgets it. If it crashes, any other instance — even one that has just been started — can take over within one second and finish what the previous one started. In our system, the coordinator is not even a separate server. It is built into every background process. That is why I say that there is no central coordinator in the system. There is no single point whose failure could stop transactions.

Another important property of Phantom is that if a transaction affects only one partition, the entire question-and-answer process is skipped. The unit verifies, writes, and commits in a single round. That roughly halves the execution time. And because revenue-type accounts can be booked on any unit, we can execute most transactions while touching only a single partition. This is where the two decisions — account type and protocol design — come together, and where the figure of sub-1 ms execution for certain operations comes from.

The most beautiful idea in Phantom, for me, is not technical. The industry has spent decades making coordinators fault-tolerant. We said: make it disposable instead. If there is nothing inside it worth protecting, then its failure simply isn’t an event. We intend to open-source the protocol in the future — not the BlockBen-specific components. This could make many distributed systems simpler, faster, and more reliable.

Why Can a Ledger Run on 1 CPU?

This is a very interesting question, but I’ll try to explain it in simple terms.

The Expensive Secret of Blockchains

The most expensive thing in blockchain technology is consensus: the process through which network participants convince themselves that a change really happened and that everyone is seeing the same thing.

Public blockchains solve this by having every participant process every transaction and store the entire chain. That is why a public blockchain node can carry several terabytes of data, and why adding more machines does not necessarily make the network faster. Instead, it tends to be constrained by the slowest participants.

This is why blockchains have a reputation for being slow and resource-hungry.

Proof Instead of Copies

In the Qantrum system, trust does not come from everyone seeing everything. It comes from cryptographic proofs. That changes the entire equation.

We do not need — and should not need — to send a multi-terabyte dataset to everyone. It is enough to store a single, constant-size proof for each Ledger that can authenticate all transactions processed up to that point.

Yes, you read that correctly: the proof of the blockchain’s complete state is the same size whether it contains one hundred transactions or one hundred million. Think of it like a seal on the final page of a book that authenticates the entire book. If anyone changes even a single letter on any previous page, the seal no longer matches.

We do not use a general-purpose, all-knowing proving system of the kind you may have read about in the industry, many of which are expensive and slow. Instead, we developed purpose-built protocols for the handful of specific claims that a financial ledger needs to prove. That is why proof generation is extremely fast. And that is why a Ledger only needs 1 CPU and 2–3 GB of memory.

If we want to make it faster, we do not buy a bigger machine. We divide the blockchain state into multiple parts, let multiple Ledgers work on them, and build multiple blockchains in parallel.

Multiple Chains Witnessing Each Other

A fair question follows: if multiple blockchains are being built in parallel, doesn’t the system become a collection of disconnected chains that know nothing about each other? How can the entire system be verified from a single point?

The solution is simple. At regular intervals, every Ledger writes its own seal into the other Ledger chains, just like any other transaction. In the other chain, this becomes a normal entry, with its own proof and its own seal. Over time, every chain contains a reference to the state of every other chain.

It is like two ledgers regularly signing each other’s last page. If someone falsifies one of them, the other exposes it.

There is no central blockchain. There is a network in which every chain witnesses the others — and any transaction can be verified starting from any of them. This is the layer used by the QantrumScan verifier when you enter the Signed Data Hash.

What Happens If Something Fails?

Currently, two Ledgers — two independent blockchain-producing units — are running in the BlockBen system, each consisting of a group of three machines. The same architecture is used for the balance management layer. The three machines elect a leader. The leader does the work, while the other two monitor it, and if the leader fails, it is immediately replaced. All three machines run in separate locations, independently of one another.

The current infrastructure is sized for approximately 2,000 transactions per second. If we need more capacity, we do not migrate anything. We simply open another unit and redirect part of the workload to it — while the system is running, without downtime — and capacity increases with the number of units.

But it is not enough to simply write this down. It has to be tested. And not just because we think it works: the DORA European regulation requires financial service providers to regularly test how their systems behave when something goes wrong.

If one machine in a group fails: nothing happens. A new leader takes over very quickly, almost unnoticed, and processing continues. No rejected transactions.

If an entire group fails: that particular unit stops. The other unit remains operational — this is what we call partial operation. The affected customers are temporarily unable to initiate transactions, while everyone else can continue. Recovery is completely automatic: once the machines become available again, they pick up the work and continue processing.

During testing, we examined server failures, network route failures, and failures of other system components. There was also a monkey test: we did not trigger the failures ourselves. Instead, a program specifically written for this purpose generated random failures in the infrastructure while the system was under load.

The goal is not to create a system that never fails. Such a system does not exist. The goal is to make sure that every failure has a defined response, and that no single component failure leaves the system in a state from which it cannot recover automatically.

The Numbers — and What They Actually Mean

  • ~30–40× faster blockchain and balance management compared to the previous version.
  • Sub-1 millisecond execution time for certain operations.

These are internal measurements. They measure how long it takes for a decision to be made and finalized at the heart of the system. They do not include the journey between the mobile application and the server, login, verification, or other checks — in other words, they do not measure how quickly the result appears on your screen.

We consider this number important because it is the part that depends on the architecture, and the part that matters as the load increases. What you see on the screen has also improved, but that is no longer determined by this layer.

The 30–40× improvement does not come from a single trick. It comes from the fact that most transactions can be completed in a single round. It comes from balance checks running in memory. And it comes from the fact that the two components now run on a common foundation written in Rust, which we designed properly once instead of maintaining two separate implementations.

What’s Next?

The smart contract executor is still running on the original version. That is the next step. Our first measurements on the test version of the new executor show 40–50× acceleration — but it is still in the testing phase, the numbers may change, and production is always different. When it goes live, we will report the real-world results in the same way.

There is no timeline yet for the open-source release of the Phantom protocol. The documentation is ready, and the system is already running on it in production, but releasing it requires documentation, a reference implementation, and a test suite. We do not want to publish something half-finished.

The current two units are sufficient for today’s and the foreseeable workload. We have not yet needed to open a new unit in production, although we have already done so in testing. When we need to do it in production for the first time, that will be a separate article as well. 😉

The Toaster Case — or Why You Have to Be Disruptive

So, what exactly is the story behind the toaster? 😃

Half a year ago, during a meeting, when we first defined these goals, one of our developers asked me — I can still hear his calm voice in my head:

“Do you really want the blockchain producer to run on a toaster? You’re crazy…”

It was a fair question.

The common perception of blockchain systems is that they are resource-hungry, slow, and difficult to scale. That is often true. But it is not because blockchain has to be that way. It is because most systems build trust by having everyone copy everything.

If trust comes from proofs rather than copies, a machine only needs to organize and verify its own part. And that can run on 1 CPU.

We are building financial infrastructure. The expectations are different there. And, as it turns out, so are the possibilities.

Looks like we weren’t completely crazy after all. 😃

We imagined it. We built it. And now it works.

And Why Do We Need This Much Performance?

I think you can probably guess by now.

With the launch of StockLock, we expect a significant increase in transaction volume from our existing customers. Future services will use the same foundation, and we expect new users to join as well.

We did not size the system for today’s traffic. We sized it for what comes next — and built it so that when that demand arrives, we do not have to migrate anything. We simply switch on another processing unit.

That is why we rewrote it. That is why there is no central coordinator. That is why it can scale without downtime. That is why it runs on a toaster.

We imagined it. We built it. And now it works.

#BlockBen #StockLock #Qantrum