Concurrency is the ability of a system to make progress on multiple tasks during overlapping periods of time.
Whenever we talk about concurrency, you will also hear people talk about parallelism. The two ideas are related, but they are different.
A concurrent program can have several tasks in progress even if only one task is executing instructions at a particular point in time. The system can execute one task for a while, stop it, execute another task, and later return to the first one.
Parallelism goes further. With parallelism, multiple tasks are actually executing at the same time. This requires multiple execution units, such as multiple CPU cores.
Think about a juggler.
A single juggler can keep several balls in motion. The juggler is still one person. At any particular instant, the juggler may be interacting with only one or two balls, but all the balls are part of the same activity and they are all making progress over time.
That gives us a simple picture of concurrency.
ball A
*
/ \
/ \
* *
ball B ball C
O
/|\
/ \
juggler
Now imagine two jugglers standing beside each other, with each person juggling their own set of balls.
Both jugglers can perform work at exactly the same time.
That gives us a picture of parallelism.
Juggler A Juggler B
* *
* * * *
O O
/|\ /|\
/ \ / \
A program can therefore be concurrent without being parallel.
A single-core machine can run several programs concurrently. The operating system switches between them quickly enough that they all appear to be making progress.
A multi-core machine can also run some of those programs in parallel because different cores can execute instructions at the same time.
This difference will become important later because Node.js, Apache Tomcat, and Tokio all provide concurrency, but they do not represent concurrent work in exactly the same way.
They also do not use CPU cores in exactly the same way.
Before we get to them, we first need to understand the problem concurrency is trying to solve.
Why Concurrency?
If you are coming across concurrency for the first time, you may have a very reasonable question:
Why do we need it?
Why can’t a program simply finish one task before starting another one?
The easiest answer is that programs spend a lot of time waiting.
Consider a web server handling a request.
The request arrives and the application needs information from a database.
HTTP request
↓
application
↓
database query
The application sends the database query.
At this point, the CPU does not need to continuously execute instructions for that request.
The query has already been sent.
The database may be running on another machine. It needs to receive the query, find the requested data, create a response, send that response through the network, and eventually the operating system on our machine receives those bytes.
That takes time.
From the application’s point of view, a large part of that time is simply waiting.
We can imagine the request like this:
CPU work
↓
send database query
↓
waiting...
waiting...
waiting...
↓
database response arrives
↓
CPU work
↓
send HTTP response
If our server insisted on completing this entire request before starting another one, it would spend a lot of useful time doing nothing.
Imagine request A starts first.
Request A
↓
query database
↓
wait 80ms
↓
continue
Now request B arrives while request A is waiting.
With a completely sequential model, request B cannot start yet.
Request A: work ───────── waiting ───────── work
Request B: work ─ waiting ─ work
The CPU may be available during much of request A’s waiting period, but request B is still forced to wait.
Concurrency allows us to use that time.
Request A: work ───────── waiting ───────── work
|
Request B: work ─── waiting ─── work
|
Request C: work ─── waiting ─── work
While request A is waiting for the database, request B can begin.
When request B reaches something it needs to wait for, request C can make progress.
Eventually the database response for request A arrives and request A can continue.
This is why concurrency matters so much in network servers.
A typical web request may need to:
receive HTTP request
↓
query PostgreSQL
↓
read from Redis
↓
call another HTTP service
↓
publish a message
↓
send HTTP response
Some parts require CPU time.
Many other parts involve waiting for another system.
Concurrency gives us a way to keep many of these operations in progress without requiring one request to completely finish before the next request can start.
This can increase the number of requests a server can handle within a period of time.
There are many details hidden inside that simple statement, though.
If several pieces of work are in progress, something has to keep track of them.
Something has to know which work can currently execute.
Something has to know which work is waiting.
Something has to decide when waiting work is allowed to continue.
Something also has to preserve enough information for that work to continue from where it stopped.
This is where the different models of concurrency begin to appear.
The Traditional Unit of Concurrency: The OS Thread
For a long time, one of the main ways programmers have represented concurrent work is through operating system threads.
A process can contain several threads.
For example:
Process
│
├── Thread 1
│
├── Thread 2
│
├── Thread 3
│
└── Thread 4
These threads share resources that belong to the process, such as its address space, but each thread has its own execution context.
You can think of an execution context as everything required for the operating system to stop a thread and later allow it to continue from the same place.
This includes things such as its stack, CPU register state, current instruction position, and scheduling information.
Suppose we have two threads.
Thread A
functionOne()
↓
functionTwo()
↓
database.query()
↓
waiting
Thread B
handleRequest()
↓
calculateSomething()
↓
continue...
When Thread A reaches a blocking database operation, the operating system can put it to sleep.
Thread B can then execute.
Later, when the operation Thread A was waiting for is ready to continue, the operating system can schedule Thread A again.
Thread A still has its stack.
It still knows which functions called which functions.
Its local variables are still available.
Its execution context has been preserved.
From the programmer’s point of view, this makes blocking code very convenient to write.
For example:
user = database.getUser(userId)
account = database.getAccount(user.accountId)
return createResponse(user, account)
The code reads from top to bottom.
getUser() appears to return a value.
Then getAccount() runs.
If getUser() needs to wait for the network, the thread itself can wait. When the operation finishes, the thread continues from the same point.
This gives us an important mental model that we will return to throughout this article:
Thread-based concurrency:
"Preserve my execution context while I wait."
Creating a thread
In C, POSIX threads give us a fairly direct way to see this model.
A small example looks like this:
#include <stdio.h>
#include <pthread.h>
#include <stdlib.h>
typedef struct {
int a;
int b;
} myarg_t;
typedef struct {
int x;
int y;
} myret_t;
void *mythread(void *arg) {
myarg_t *args = (myarg_t *) arg;
printf("received %d %d\n", args->a, args->b);
myret_t *result = malloc(sizeof(myret_t));
result->x = 1;
result->y = 2;
return result;
}
int main(void) {
pthread_t thread;
myarg_t args;
args.a = 10;
args.b = 20;
int rc = pthread_create(
&thread,
NULL,
mythread,
&args
);
if (rc != 0) {
fprintf(stderr, "failed to create thread\n");
return 1;
}
myret_t *result;
pthread_join(thread, (void **) &result);
printf("returned %d %d\n", result->x, result->y);
free(result);
return 0;
}
The main thread creates another thread.
After that happens, there are two execution contexts in the process.
Conceptually:
Process
│
├── Main Thread
│ │
│ ├── pthread_create()
│ │
│ └── pthread_join()
│
└── New Thread
│
└── mythread()
The operating system can schedule these threads independently.
If the machine has one available CPU core, the operating system can switch between them.
CPU:
Thread A
↓
Thread B
↓
Thread A
↓
Thread B
They are running concurrently.
If the machine has multiple available CPU cores, Thread A and Thread B may also execute at the same time.
CPU Core 1: Thread A
CPU Core 2: Thread B
Now we have both concurrency and parallelism.
Thread pools
A web server generally does not want to create a brand new operating system thread for every tiny piece of work and then immediately destroy it.
Creating and destroying threads has a cost.
Instead, servers commonly maintain a collection of reusable worker threads.
This is a thread pool.
┌── Worker Thread 1
│
incoming requests ├── Worker Thread 2
↓ │
request queue ────────┼── Worker Thread 3
│
├── Worker Thread 4
│
└── Worker Thread 5
A request enters a queue.
An available worker thread takes the request.
That thread runs the application code.
If the application performs blocking I/O, the worker thread may sleep while it waits.
Another worker thread can continue processing another request during that time.
This model has been used successfully for a very long time and is still used by a huge number of production systems.
It also gives programmers a simple way to reason about request processing.
One request can often look like one normal sequence of function calls:
receive request
↓
call service
↓
query database
↓
process response
↓
send result
The complexity of suspending and resuming the thread is handled by the operating system.
There is still a cost to this model.
Every thread needs its own execution state.
Each thread needs stack space.
The operating system scheduler needs to manage the threads.
When the CPU switches from one thread to another, it performs a context switch.
With a moderate number of threads, these costs may be perfectly acceptable.
Things become more interesting when the number of concurrent operations becomes very large.
Imagine a server with:
100 connections
Then:
1,000 connections
Then:
10,000 connections
Then:
100,000 mostly idle connections
If every connection needs its own thread sitting around while it waits for network activity, the amount of state being preserved begins to matter.
This raises an important question.
Do we always need to preserve an entire thread just because some network operation is waiting?
That question leads us into event-based concurrency.
Waiting Is Where Concurrency Becomes Interesting
It is easy to picture concurrency when several tasks are actively using the CPU.
The more interesting part is what happens when most of those tasks are doing nothing.
Consider a web application that needs to retrieve a user from a database.
The code might conceptually look like this:
handleRequest()
↓
query database
↓
process user
↓
build response
Imagine that the total request takes 100 milliseconds.
It is very unlikely that the CPU spent all 100 milliseconds continuously executing instructions for that request.
Maybe the actual work looks more like this:
0ms 2ms 92ms 95ms 100ms
|----------|----------------------------|-------|----------|
CPU work waiting for database CPU send
work response
The request existed for 100 milliseconds.
The CPU may have spent only a small amount of that time actually working on it.
The rest of the time was spent waiting for something outside the current computation.
The same thing happens with network calls.
application
↓
send request to payment provider
↓
wait
↓
wait
↓
wait
↓
response arrives
↓
continue
It happens with Redis.
It happens with files.
It happens when waiting for another microservice.
It happens when waiting for a client to send more bytes over a TCP connection.
For servers that handle a lot of I/O, waiting can make up a large part of the lifetime of a request.
Now imagine 10,000 connections.
Most of them may not need the CPU right now.
They simply need the system to remember:
Connection 17 is waiting for bytes.
Connection 308 is waiting until it can write more data.
Connection 901 is waiting for a database result.
Connection 4,212 is waiting for an HTTP response.
Connection 8,991 currently has data available.
At that point the main problem is no longer:
How do I execute 10,000 things at the same time?
The machine probably does not have 10,000 CPU cores anyway.
The more useful question becomes:
How do I keep track of 10,000 operations that may be in different stages of progress, while only spending CPU time on the operations that can actually do something right now?
That is a very different problem.
A thread-based system has one answer.
It can associate execution contexts with the work and allow waiting threads to sleep.
An event-based system uses another representation.
Instead of keeping a thread tied to each waiting operation, it can keep the state required to continue the operation later.
This gives us the second mental model:
Event-based concurrency:
"Preserve the state I need,
and continue me when I can make progress again."
This difference is at the heart of Node.js.
It is also at the heart of asynchronous Rust.
We will eventually see that Apache Tomcat also makes use of some of the same underlying operating system ideas, even though the programming model exposed to application code can look very different.
Event-Based Concurrency Is Still Concurrency
When many programmers first learn concurrency, they learn it through threads.
That creates a very strong mental connection:
concurrency = multiple threads
It is an understandable connection.
Threads are one of the most visible ways to represent concurrent execution.
But threads are not the definition of concurrency.
Remember the definition we started with.
Concurrency means that multiple tasks can remain in progress during overlapping periods of time.
Consider three network operations.
Operation A
Operation B
Operation C
We start Operation A.
It sends a network request and cannot continue until a response arrives.
Instead of sitting there and waiting, we start Operation B.
Operation B also reaches a network operation and has to wait.
Then we start Operation C.
Our system now looks something like this:
Operation A: started ───────── waiting ────────────── continue
Operation B: started ───────── waiting ─── continue
Operation C: started ───── waiting ───────── continue
All three operations are in progress.
They overlap in time.
They are concurrent.
There may still be only one thread executing application code at any particular instant.
That does not change the fact that several operations are in progress.
This is the idea behind event-based concurrency.
A simplified event-driven system might behave like this:
event loop
↓
check for work that can make progress
↓
run a handler
↓
handler starts asynchronous operation
↓
handler cannot continue yet
↓
save required state
↓
return to event loop
↓
run some other ready work
Later, something happens.
Maybe bytes arrive from the network.
Maybe a socket becomes writable.
Maybe a timer expires.
Maybe another operation finishes.
That event tells the system that some previously suspended work may now be able to make progress.
The event loop can then return to that work.
Operation A
↓
start network read
↓
cannot continue
↓
return control
Operation B
↓
do some work
↓
return control
Operation C
↓
do some work
↓
return control
network event for A arrives
↓
continue Operation A
Notice what has changed.
With the thread model, the waiting operation can keep an execution context around.
Thread model:
"Keep this execution context for me.
I will continue from here when I wake up."
With the event model, the system keeps enough information to continue the operation later.
Event model:
"I cannot make progress now.
Remember what I need and run me again
when the event I am waiting for happens."
The work has not disappeared.
It has simply stopped using an execution resource while it waits.
One thread can therefore manage many concurrent operations
Suppose we have five TCP connections.
Connection A
Connection B
Connection C
Connection D
Connection E
At this particular moment:
A: waiting for data
B: waiting for data
C: data available
D: waiting for data
E: waiting until socket is writable
Only Connection C may have useful work for the CPU right now.
An event-driven system does not need to continuously execute code for A, B, D, and E.
It can run the handler for C.
A waiting
B waiting
\
\
Event Loop
|
└── C ready → run C
/
/
D waiting
E waiting
After C makes as much progress as it can, control returns to the event loop.
The system waits until something else becomes ready.
This is still concurrency.
The operations overlap in time.
They have independent state.
They can become ready at different times.
They can finish in a different order from the order in which they started.
The main difference is how the waiting work is represented.
Event-based does not mean single-core
There is another misconception worth removing early.
Event-based concurrency does not mean that the entire application must run on one CPU core.
You can have several event loops.
Those event loops can run on different operating system threads.
Those threads can run on different CPU cores.
Conceptually:
CPU Core 1
↓
Thread 1
↓
Event Loop 1
↓
many tasks
CPU Core 2
↓
Thread 2
↓
Event Loop 2
↓
many tasks
Now the system has event-based concurrency and CPU parallelism at the same time.
This idea will become very important when we get to Tokio.
Tokio can manage a very large number of asynchronous tasks while also using several worker threads.
Those worker threads can execute on several CPU cores.
Node.js can also use several threads and processes even though JavaScript execution is commonly introduced through the idea of a single event loop.
Real systems frequently combine these models.
They use events in some places.
They use threads in other places.
They use thread pools when blocking work cannot easily be represented as non-blocking I/O.
They may also use several processes when they want stronger isolation or access to more CPU cores.
Before looking at those higher-level systems, we need to go one level lower.
There is an obvious question.
How does an event loop know that a network operation can make progress?
The operating system has to help.
How the Operating System Makes Event-Based Concurrency Possible
Let us start with a normal TCP connection.
Our program has a socket.
It wants to read data from that socket.
With a blocking operation, the program can say, conceptually:
read(socket)
If no data is available, the calling thread can sleep until the operating system has something to return.
The programming model is straightforward.
call read()
↓
data available?
↓
no
↓
thread sleeps
↓
data arrives
↓
thread wakes up
↓
read() returns
The operating system does the waiting for us.
That works very well.
The problem appears when we want one thread to manage a large number of sockets without blocking on any single one of them.
Imagine we have:
socket 1
socket 2
socket 3
socket 4
...
socket 10,000
If our thread calls a blocking read() on socket 1 and socket 1 has no data, the thread goes to sleep.
Maybe socket 8,721 already has data waiting.
Our thread cannot get there because it is sleeping on socket 1.
One answer would be to assign different threads to different sockets.
Another answer is to make the sockets non-blocking and ask the operating system to help us keep track of readiness.
Non-blocking I/O
With a non-blocking socket, the application can attempt an operation without agreeing to sleep if the operation cannot make progress immediately.
Conceptually:
read(socket)
↓
data available?
/ \
yes no
| |
read return "try again later"
data
This solves one problem.
Our thread no longer has to sleep just because one socket has no data.
But it creates another problem.
How do we know when to try again?
We could repeatedly check every socket.
check socket 1
check socket 2
check socket 3
check socket 4
...
check socket 10,000
start again
check socket 1
check socket 2
...
That would waste a lot of CPU time.
Most sockets may have nothing happening.
We need the operating system to tell us which sockets are interesting right now.
Operating systems provide interfaces for this.
On Linux, one of the important interfaces is epoll.
BSD systems and macOS provide kqueue.
Windows has I/O Completion Ports, usually called IOCP. The Windows model is based more directly around completion notification, so it is not identical to readiness systems such as epoll, but it solves the same broad problem of efficiently managing large numbers of asynchronous I/O operations.
For the Linux example, we can think about epoll like this:
application
↓
"Please watch these sockets"
socket 1
socket 2
socket 3
socket 4
...
socket 10,000
↓
kernel
The application does not need to continuously scan all 10,000 sockets itself.
Instead, it can wait for the kernel to report sockets that are ready for some kind of I/O progress.
For example:
10,000 registered sockets
↓
epoll
↓
kernel reports:
socket 17 readable
socket 912 readable
socket 4001 writable
The application can now focus on those sockets.
event loop
↓
ask kernel for ready file descriptors
↓
socket 17 is ready
socket 912 is ready
socket 4001 is ready
↓
run code associated with those sockets
There is an important detail here.
Readiness does not necessarily mean:
Your entire operation has completed.
It means something closer to:
This resource may now be able to make some progress.
The application still performs the actual read or write operation.
Suppose socket 17 has some bytes available.
The system reports it as readable.
Our event loop receives that information.
epoll
↓
socket 17 readable
↓
event loop
↓
read(socket 17)
↓
process bytes
After processing those bytes, the application may discover that it needs more data.
It can once again wait for another readiness notification.
This is the operating system foundation behind a large part of event-driven network programming.
The event loop is not constantly spinning
The phrase “event loop” sometimes gives people the wrong picture.
They imagine code like this:
while (true) {
checkEverything();
}
where the CPU continuously runs in circles checking whether something has happened.
A real event loop can sleep inside an operating system call when there is no useful work to perform.
Conceptually:
event loop
↓
anything ready to run?
↓
no
↓
ask operating system to wait
↓
thread sleeps efficiently
↓
network activity occurs
↓
kernel wakes thread
↓
event loop receives ready events
↓
process them
This allows one execution thread to manage large numbers of mostly waiting network connections.
The thread consumes CPU time when useful work is available.
When nothing can make progress, it can sleep.
We now have two ways to represent waiting
At this point we can put the two models beside each other.
The thread-oriented model might look like this:
Connection A → Thread A → blocked
Connection B → Thread B → running
Connection C → Thread C → blocked
Connection D → Thread D → blocked
The operating system preserves each thread’s execution context.
An event-oriented model can look more like this:
Connection A ─┐
Connection B ─┤
Connection C ─┼──→ OS I/O readiness
Connection D ─┤
Connection E ─┘
↓
event loop
↓
ready handlers
The program keeps state for the connections.
The operating system tells the program when some of those connections may be able to make progress.
The program then runs the corresponding code.
This gives us the distinction we will need for the rest of the article:
Thread-oriented concurrency:
Preserve an execution context while work waits.
Event-oriented concurrency:
Preserve the state required to continue the work, then schedule it again when an event says that progress may be possible.
Neither idea removes the need for the operating system.
Neither idea removes the need for threads entirely.
Neither idea automatically gives us parallelism.
They are different ways of organising concurrent work.
And once we understand that, Node.js, Apache Tomcat, and Tokio start to look much less like three unrelated technologies.
They are all dealing with the same underlying questions:
-
How do we represent work that is currently in progress?
-
What do we do when that work has to wait?
-
How do we know when it can continue?
-
Where do we store the information required to continue it?
-
And when it becomes runnable again, what execution resource should run it?
Those questions will take us directly into event loops, callbacks, futures, async/await, thread pools, and eventually the different choices made by Node.js, Tomcat, and Tokio.
Event-Based Concurrency Moves Some Work Into Application State
There is one part of the event model that we have avoided so far.
If we stop preserving an entire execution context while an operation waits, where does everything that the operation needs go?
Consider a normal blocking function:
function handleRequest(request):
user = getUser(request.userId)
account = getAccount(user.accountId)
response = buildResponse(user, account)
return response
There is useful state here.
We have request.
Later we have user.
Later we have account.
The function also has a current position in the program.
When getUser() returns, execution should continue with getAccount().
When getAccount() returns, execution should continue with buildResponse().
With a thread-oriented blocking model, the thread’s stack naturally keeps track of much of this for us.
The function can stop inside getUser() and continue later.
The call stack is still there.
Now imagine that getUser() works with non-blocking I/O.
It sends a request to another machine and then discovers that no response is available yet.
We do not want the event-loop thread to sit there and wait.
So handleRequest() needs to give control back to the event loop.
The problem is that we still need all of this information later:
request
userId
where execution stopped
what should happen when the result arrives
When the response finally arrives, something has to say:
We were processing this request. We had reached this particular point. Here is the data that was returned. Continue from here.
Early event-driven programs made this relationship very visible.
The programmer would start an operation and provide a function to execute when that operation completed.
Conceptually:
startGetUser(userId, whenUserIsReady)
Then:
function whenUserIsReady(user):
startGetAccount(user.accountId, whenAccountIsReady)
Then:
function whenAccountIsReady(account):
response = buildResponse(account)
sendResponse(response)
The continuation of the program is now represented by functions.
Instead of having one execution context paused halfway through the work, we split the work into pieces.
Part 1
↓
start I/O
↓
return to event loop
event happens
↓
Part 2
↓
start another I/O operation
↓
return to event loop
event happens
↓
Part 3
This is where the word continuation becomes useful.
A continuation is simply a way of describing what should happen next.
For example:
database query finishes
↓
what happens next?
↓
processUser()
processUser() is part of the continuation of that computation.
You do not need to know continuation-passing style or programming language theory to understand the important idea here.
The event-based model needs a way to remember two things:
1. the data required to continue
2. what code should run next
We have moved some of the responsibility that was naturally represented by the thread’s execution context into another form.
This is the trade that makes event-driven systems possible.
Callbacks make the continuation explicit
JavaScript programmers who worked with Node.js before Promises and async/await became common will recognise this style immediately.
You might have written code similar to this:
getUser(userId, function (err, user) {
if (err) {
handleError(err);
return;
}
getAccount(user.accountId, function (err, account) {
if (err) {
handleError(err);
return;
}
buildResponse(user, account, function (err, response) {
if (err) {
handleError(err);
return;
}
sendResponse(response);
});
});
});
There is nothing mysterious happening here.
Each callback is describing what should happen when an earlier operation reaches the point where progress can continue.
The event loop does not need to keep the original JavaScript function sitting on the stack waiting for every operation to finish.
Instead, the required state is captured and the callback can run later.
The problem becomes obvious as the number of steps increases.
Our simple request can turn into:
get user
↓
callback
↓
get account
↓
callback
↓
get permissions
↓
callback
↓
call another service
↓
callback
↓
write response
Error handling gets spread across these pieces.
The control flow becomes harder to read.
The programmer is thinking about one logical operation:
handle this HTTP request
but the code is split into several separate functions that execute at different times.
The machine is comfortable with this representation.
Humans usually prefer something that reads from top to bottom.
That is where Futures, Promises, and eventually async/await become useful.
Futures, Promises and async/await
Suppose an operation cannot produce its result immediately.
One way to represent that is to return an object that stands for a value that may exist later.
In JavaScript, this idea appears as a Promise.
In Rust, it appears through the Future trait.
The details are different between the languages, but the broad idea is similar.
Instead of saying:
Here is the result.
the operation can effectively say:
Here is something that represents
a result that may become available later.
This gives the program an object that can represent unfinished work.
That is useful because unfinished work is exactly what we have been talking about throughout this article.
Operation A
↓
started
↓
waiting
↓
unfinished
The operation still exists.
It simply cannot produce its final value yet.
Why async/await matters
Now consider this JavaScript:
async function handleRequest(userId) {
const user = await getUser(userId);
const account = await getAccount(user.accountId);
return buildResponse(user, account);
}
Compare that to the callback version:
getUser(userId, function (err, user) {
getAccount(user.accountId, function (err, account) {
buildResponse(user, account);
});
});
The first version looks much closer to normal sequential code.
We can read it from top to bottom.
get user
↓
get account
↓
build response
This is where async/await can create a misleading mental model if we do not understand what we have already discussed.
When we write:
const user = await getUser(userId);
it looks like the whole execution thread must stop there.
That is not what asynchronous await means.
The current asynchronous computation reaches a point where it cannot continue until another operation produces a result.
The runtime can allow other work to execute.
When the result becomes available, the computation can continue from the point after the await.
Conceptually:
handleRequest()
↓
getUser()
↓
await
↓
result ready?
/ \
yes no
| |
continue suspend this
computation
↓
run other work
↓
result becomes ready
↓
resume computation
This should look familiar.
We are still doing what we described in the event model earlier:
"I cannot make progress now.
Keep the state required to continue me.
Run something else.
Come back to me when progress is possible."
async/await gives us a much nicer way to write that program.
It does not remove the event-driven model underneath it.
The syntax changed. The underlying problem did not.
This is important enough to spend some time on.
Consider:
const response = await fetch(url);
and:
let response = client.get(url).send().await?;
The syntax in both cases gives us code that looks sequential.
The machine still has to answer all of the same questions we have been asking since the beginning:
What state must be preserved?
How do we know the operation is waiting?
How do we know it can continue?
Who schedules it again?
Where does it execute when it resumes?
Those questions do not disappear because we wrote .await.
The language and runtime are helping us manage them.
This is one of the main ideas I want us to carry into the rest of the article:
async/awaitis a way of writing asynchronous control flow. It does not turn asynchronous work into blocking work.
We can now look at Node.js, Tomcat, and Tokio with a much stronger mental model.
Node.js: The Event Loop Is Only Part of the Story
Node.js is probably the system most people associate with event-based concurrency.
It is also commonly explained with one sentence:
Node.js is single-threaded.
That sentence is useful for introducing one part of Node.js, but it leaves out enough details that it can eventually make the runtime harder to understand.
Let us start from the part that is true.
In a normal Node.js program, JavaScript callbacks for a given Node.js environment are executed through an event loop.
Imagine an HTTP server:
import http from "node:http";
const server = http.createServer((req, res) => {
res.end("Hello");
});
server.listen(3000);
When multiple requests arrive, Node.js does not normally create one JavaScript execution thread for every request.
Instead, we can think about the high-level flow like this:
client connections
↓
operating system
↓
libuv
↓
event loop
↓
JavaScript callback
libuv is the library underneath Node.js that provides the event loop and cross-platform asynchronous I/O facilities.
On Linux, network I/O can use operating-system readiness mechanisms such as the epoll model we discussed earlier. On other operating systems, libuv uses the corresponding facilities available there.
This means Node.js can have a very large number of network operations in progress while only a small number of threads are involved.
The official Node.js documentation describes the runtime as using an event loop for JavaScript callbacks and non-blocking I/O, along with a worker pool for operations that need a different execution model.
Let us make that more concrete.
Imagine three requests:
Request A
Request B
Request C
Request A starts and needs to call PostgreSQL.
Request A
↓
send database query
↓
await response
Once that network operation has been started, there is no reason for the JavaScript execution thread to sit idle waiting for PostgreSQL.
Request B can execute.
Request A: waiting for PostgreSQL
Request B: executing JavaScript
Request C: waiting for its turn
Request B might then start a Redis operation and also become suspended.
Request A: waiting for PostgreSQL
Request B: waiting for Redis
Request C: executing JavaScript
Eventually the database response associated with Request A arrives.
The operating system reports network activity.
libuv processes the event.
Node.js can then arrange for the JavaScript continuation associated with Request A to run.
PostgreSQL response
↓
operating system
↓
libuv
↓
event loop
↓
continue Request A
This is event-based concurrency.
Three requests have been in progress during overlapping periods of time.
The JavaScript code for all three requests did not need to execute at the same instant.
Where await fits into Node.js
Suppose our request handler looks like this:
async function handleRequest(req, res) {
const user = await getUser(req.userId);
const account = await getAccount(user.accountId);
res.end(JSON.stringify(account));
}
When execution reaches:
const user = await getUser(req.userId);
and the underlying operation is not finished, that asynchronous function can suspend.
Other JavaScript work can run.
When the Promise being awaited settles, the continuation of the async function becomes eligible to run again.
So the code looks like this:
handleRequest()
↓
getUser()
↓
await
↓
suspend
Meanwhile:
event loop
↓
run other callbacks
↓
process other requests
↓
handle timers
↓
process other ready work
Later:
getUser() result becomes available
↓
Promise settles
↓
continuation becomes runnable
↓
handleRequest() continues
We have once again reached the same model:
preserve the state required to continue
+
run the continuation later
The syntax has made this much easier to write than the deeply nested callback version.
Node.js also has a worker pool
Now we reach the part that gets lost when Node.js is described only as “single-threaded”.
Not every asynchronous operation can be handled through network readiness.
Consider filesystem operations.
Depending on the operating system and the operation involved, there may not be a convenient readiness interface that works like a TCP socket.
Node.js also has operations such as some DNS functions, cryptographic work, and compression that can require blocking or CPU-intensive work.
libuv maintains a worker pool for this type of work. Node.js currently documents filesystem APIs, dns.lookup(), several crypto APIs, and zlib APIs among the operations that make use of this pool.
Conceptually:
Node.js
Event Loop
/ \
/ \
network I/O work that needs
| worker threads
| |
↓ ↓
OS readiness libuv worker pool
Suppose our JavaScript does this:
const content = await fs.promises.readFile("large-file.txt");
From the JavaScript programmer’s point of view, this is still asynchronous.
The JavaScript execution thread does not need to remain blocked until the read finishes.
Underneath the API, work may be performed through libuv’s worker pool.
When the operation completes, its result eventually makes its way back to the event loop and the associated asynchronous continuation can run.
So Node.js is already a hybrid concurrency system.
It uses operating-system event mechanisms where they fit.
It uses worker threads where they fit.
It gives JavaScript programmers one mostly consistent asynchronous programming model over those different mechanisms.
That is an important idea.
An asynchronous API does not automatically tell us how the operation is implemented underneath.
This:
await something();
could eventually depend on:
non-blocking network I/O
or:
work running in another thread
or:
another process
or even:
a value that is already available
The await tells us how our computation behaves while waiting for the result. It does not fully describe how the result is produced.
What happens when JavaScript itself takes too long?
Now imagine this code:
app.get("/calculate", (req, res) => {
let total = 0;
for (let i = 0; i < 10_000_000_000; i++) {
total += i;
}
res.json({ total });
});
There is no I/O operation here that gives the event loop an opportunity to work on another request.
The JavaScript callback keeps executing.
event loop
↓
Request A callback
↓
huge CPU calculation
↓
still calculating
↓
still calculating
↓
finish
↓
event loop gets control back
During that period, other JavaScript callbacks associated with that event loop do not get their turn.
The Node.js documentation explicitly warns about this. A long-running callback blocks the event loop from serving other clients.
This is one of the costs of having a small number of execution threads responsible for large numbers of concurrent operations.
Each piece of work needs to give control back in a reasonable amount of time.
For CPU-heavy JavaScript work, Node.js provides worker_threads, which allow JavaScript to execute on additional threads. Node.js recommends worker threads mainly for CPU-intensive JavaScript rather than ordinary asynchronous I/O.
A Node.js application can therefore look like this:
Main Node.js thread
|
Event Loop
|
┌────────────────┼────────────────┐
│ │ │
Request A Request B Request C
|
CPU-heavy work
|
↓
Worker Thread
Node.js can also use multiple processes through facilities such as cluster or normal child processes. The cluster module can run several Node.js worker processes and distribute network connections among them.
So when someone says:
Node.js is single-threaded.
A more useful mental model is:
A normal Node.js JavaScript environment executes JavaScript through an event loop, while Node.js and libuv use other threads and operating-system mechanisms underneath when required.
That gives us a much better foundation for understanding how Node.js achieves concurrency.
And it prepares us for Apache Tomcat, which makes a very different choice at the application layer.
Apache Tomcat: Threads at the Application Layer, Events at the Network Layer
Apache Tomcat gives us an interesting comparison because the programming model most Java web developers associate with Tomcat is thread-oriented.
Imagine a normal Servlet:
protected void doGet(
HttpServletRequest request,
HttpServletResponse response
) throws IOException {
User user = userService.getUser(request.getParameter("id"));
Account account = accountService.getAccount(user.getAccountId());
response.getWriter().write(account.toJson());
}
This code looks completely sequential.
get user
↓
get account
↓
write response
Suppose getUser() performs a blocking JDBC query.
The thread handling this request can wait for that query.
Another request can be handled by another thread.
Conceptually:
Request A → Worker Thread 1 → waiting for database
Request B → Worker Thread 2 → executing
Request C → Worker Thread 3 → waiting for another service
Request D → Worker Thread 4 → executing
The operating system scheduler manages these threads.
When Thread 1 blocks waiting for the database, the CPU can execute another runnable thread.
This is the thread-oriented model we discussed near the beginning of the article.
Current Tomcat documentation describes a normal non-asynchronous HTTP request as requiring a request-processing thread for the duration of the request. Tomcat uses a configurable thread pool for this request processing.
At first glance, this appears completely different from Node.js.
Node.js:
many requests
↓
event loop
Tomcat:
many requests
↓
thread pool
If we stop there, however, we miss something important.
A connection does not always need a request thread
HTTP connections spend time waiting too.
A client connects to the server.
It may send a request.
The server responds.
With HTTP keep-alive, the connection may stay open because another request could arrive later.
Imagine thousands of connections:
Connection 1: waiting for next request
Connection 2: currently sending request
Connection 3: waiting for next request
Connection 4: waiting for next request
Connection 5: currently sending request
...
Connection 10,000: waiting for next request
Do we really need one request-processing thread blocked on every idle keep-alive connection?
We already know another way to handle this problem.
The operating system can monitor sockets and report which ones are ready.
Tomcat’s default HTTP connector uses Java NIO. Its NIO implementation has a poller that registers sockets, waits for socket events, and hands ready sockets to request processing when appropriate. Tomcat describes this connector as supporting non-blocking polling, while normal request processing still uses request-processing threads.
Now our picture changes.
A simplified version looks like this:
TCP Connections
A B C D E F
│ │ │ │ │ │
└──────┴──────┴──────┴──────┴──────┘
↓
Java NIO / Poller
↓
socket has useful work
↓
Thread Pool
↓
Servlet execution
This is very interesting.
At the network layer, Tomcat can use event-based I/O.
At the application request-processing layer, it can give the programmer a thread-oriented model.
This is why dividing systems into simple categories such as:
Node.js = events
Tomcat = threads
is useful only as an introduction.
Real systems contain layers.
Different layers can make different concurrency choices.
Connection concurrency and request execution are separate problems
This distinction is worth making clear.
Tomcat needs to handle the lifecycle of network connections.
That includes things such as:
accept connection
wait for request bytes
read request
wait for another request on keep-alive connection
detect timeouts
write response bytes
Then it also needs to run application code:
controller / servlet
service code
database call
business logic
response construction
These two parts of the system do not have to use the same concurrency model.
We can draw the boundary like this:
NETWORK SIDE
many TCP connections
↓
readiness polling
↓
-------------------------------------------------
APPLICATION SIDE
request available
↓
worker thread
↓
servlet code
↓
response
This is one of the most useful things Tomcat teaches us about concurrency.
There is no rule saying an application has to choose event-based concurrency everywhere or thread-based concurrency everywhere.
You can multiplex thousands of connections using readiness notifications and then hand actual request processing to a bounded set of threads.
What happens when application code blocks?
Suppose our servlet does this:
User user = database.findUser(id);
and that database API is blocking.
The request-processing thread waits.
Worker Thread 7
↓
database query
↓
blocked
↓
database response
↓
continue
That particular worker thread cannot process another request while it waits.
Other worker threads remain available.
Thread 1 → Request A → blocked
Thread 2 → Request B → running
Thread 3 → Request C → running
Thread 4 → available
This is very different from blocking Node.js’s main event-loop thread.
If one Tomcat worker blocks waiting for I/O, the remaining workers can still execute other requests.
The capacity of the pool still matters.
Imagine a pool containing 200 request-processing threads.
Now imagine 200 requests all make very slow blocking calls.
200 worker threads
↓
200 slow operations
↓
all workers occupied
Request 201 now has to wait for a worker to become available.
Tomcat’s current connector documentation describes this relationship directly. maxThreads limits the number of request-processing threads when the connector uses its internal executor, and excess work may wait until a thread becomes available.
The resource that can become exhausted is different from the Node.js event-loop case, but the underlying question is familiar:
We have a finite number of execution resources. What happens when work occupies them for too long?
We will come back to that question later.
Tomcat can also do asynchronous request processing
Modern Servlet APIs also provide asynchronous processing and non-blocking I/O.
For example, a servlet can start asynchronous processing and release the original container thread instead of keeping it tied to a request for the entire lifetime of some long-running operation.
The Servlet API also supports non-blocking reads and writes through listener-based APIs. For writes, for example, the container can notify a WriteListener when data can be written without blocking.
The model begins to look familiar again:
request
↓
start asynchronous work
↓
release request thread
↓
something happens later
↓
continue processing
We are back in the world of:
state
+
events
+
continuations
Tomcat therefore gives us an important bridge between the two concurrency models.
It can use event-driven I/O for connections.
It can use normal worker threads for application request processing.
It can also expose asynchronous facilities when applications need them.
Current Tomcat versions can additionally be configured to use Java virtual threads for request processing. That introduces another interesting concurrency model, but it deserves its own discussion. For this article, we will stay focused on the traditional thread-pool model because it gives us the clearest comparison with Node.js and Tokio.
Now we can move to Tokio.
Tokio starts from a very different programming language, but by this point many of its ideas should already look familiar.
Tokio: Futures, Tasks and a Runtime
Rust itself does not provide an asynchronous runtime.
The language gives us features such as:
async fn
and:
.await
Rust’s standard library also defines the Future trait.
Something still needs to drive those futures, schedule asynchronous tasks, interact with operating-system I/O, and decide when suspended work should run again.
Tokio provides those facilities.
This distinction matters.
The async and .await syntax belongs to Rust.
The runtime that executes large numbers of asynchronous tasks can be Tokio.
Let us begin with a simple function:
async fn get_account(user_id: u64) -> Result<Account, Error> {
let user = get_user(user_id).await?;
let account = load_account(user.account_id).await?;
Ok(account)
}
To someone coming from JavaScript, this looks very familiar.
async function getAccount(userId) {
const user = await getUser(userId);
const account = await loadAccount(user.accountId);
return account;
}
The syntax is similar because both are trying to solve a similar human problem.
We want to write asynchronous programs in a form that can still be read in a mostly sequential way.
The interesting part in Rust is what this turns into conceptually.
An async computation has state
Consider this function again:
async fn get_account(user_id: u64) -> Result<Account, Error> {
let user = get_user(user_id).await?;
let account = load_account(user.account_id).await?;
Ok(account)
}
The function can be in several states.
At the beginning:
State 0
We know:
user_id
Next:
call get_user()
Then get_user() may be unfinished.
State 1
We know:
user_id
get_user operation is still pending
Next:
wait until get_user can make progress
Later the user becomes available.
State 2
We know:
user
Next:
call load_account()
Then that operation may also become pending.
State 3
We know:
user
load_account operation is pending
Next:
wait until load_account can make progress
Finally:
State 4
We know:
user
account
Next:
return account
This should remind us of something we discussed much earlier.
Event model:
Preserve the state I need, and continue me when I can make progress again.
Rust’s async machinery gives the program a structured way to represent exactly that kind of suspended computation.
An async block or async fn produces a value that implements Future.
The standard library describes a Future as an asynchronous computation that may not have finished yet. A future is actively driven through its poll method rather than automatically running itself.
The central operation looks roughly like this:
fn poll(...) -> Poll<Output>
and the important part for our mental model is that poll can produce one of two broad outcomes:
Ready(value)
or:
Pending
You can read them almost literally.
Ready(value)
"I have completed.
Here is my result."
and:
Pending
"I cannot complete right now."
This is event-based concurrency made very explicit.
A Future does not mean a thread
This point is important for engineers coming to Rust from thread-oriented environments.
Imagine that we create 10,000 futures.
That does not mean Tokio creates 10,000 operating system threads.
A future is a representation of an asynchronous computation.
It contains the state required for that computation to make progress.
We can have:
Future A: waiting for socket
Future B: waiting for timer
Future C: runnable
Future D: waiting for socket
Future E: runnable
Only the runnable work needs CPU time.
Tokio groups asynchronous execution into tasks.
Tokio describes tasks as asynchronous units of execution that can run concurrently with other tasks. Depending on runtime configuration, a spawned task may execute on the current thread or be scheduled onto another runtime worker thread.
Conceptually:
Task A
↓
Future A
Task B
↓
Future B
Task C
↓
Future C
The runtime scheduler is responsible for deciding which runnable tasks get execution time.
Now we can begin to build Tokio from pieces.
Tokio Runtime
|
┌───────────────┴───────────────┐
│ │
Scheduler I/O Driver
│ │
↓ ↓
Tasks OS I/O events
The scheduler cares about runnable tasks.
The I/O side cares about operations that are waiting for things such as sockets to become ready.
Tokio includes scheduling, I/O, timers, synchronization facilities, and other pieces needed for asynchronous applications. Tokio’s own project documentation describes Mio, one of the lower layers in its stack, as a portable interface over operating-system evented I/O APIs.
We have seen this idea before.
On Linux, the chain can conceptually look like:
TCP socket
↓
epoll / OS event mechanism
↓
Tokio I/O driver
↓
task can make progress
↓
scheduler
↓
poll task again
The names are different from Node.js.
The programming language is different.
The runtime design has many important differences.
But the underlying problem should now feel very familiar.
What actually happens at .await?
This is one of the most useful things to understand about asynchronous Rust.
Take:
let user = get_user(user_id).await?;
A common beginner interpretation is:
pause this thread here until get_user finishes
That would be normal blocking behaviour.
The asynchronous model works differently.
Conceptually, the future containing this code is being polled.
poll get_account future
↓
make progress
↓
reach get_user().await
↓
can get_user complete now?
If yes:
Ready(user)
↓
continue
If no:
Pending
↓
return control to runtime
The runtime can now execute another task.
Task A → Pending
↓
scheduler
↓
Task B → execute
↓
Task C → execute
At some point the event Task A was waiting for happens.
Maybe bytes arrive on a socket.
The runtime needs a way to know that Task A should be tried again.
Rust Futures use a Waker for this.
When a Future returns Poll::Pending, it can arrange for the task’s waker to be notified when the Future may be able to make progress. The Rust standard library’s Future documentation describes exactly this flow: poll returns Pending, a waker is stored, and the task is woken when the operation may be ready to be polled again.
So our picture becomes:
Task A is polled
↓
waiting for socket
↓
Poll::Pending
↓
runtime runs something else
↓
socket becomes ready
↓
Waker wakes Task A
↓
Task A becomes runnable
↓
scheduler eventually polls Task A again
This is a very clean expression of the event-based model.
Remember what we said much earlier:
Event model:
I cannot make progress now. Remember what I need. Come back to me when the event I care about happens.
In Rust terms, that begins to look like:
poll()
↓
Pending
↓
Waker
↓
event happens
↓
wake()
↓
poll() again
Tokio did not invent this basic idea.
The same family of ideas existed in event loops, readiness systems, callbacks, continuations, and asynchronous I/O long before Tokio existed.
Rust’s Future abstraction gives those ideas a form that works well with Rust’s language model.
Tokio then provides a runtime capable of efficiently driving large numbers of those computations.
Tokio can use multiple threads
At this point it would be easy to assume Tokio works exactly like the common Node.js event-loop model and therefore uses one thread for asynchronous tasks.
Tokio supports both a current-thread runtime and a multi-thread runtime.
The commonly used multi-thread runtime creates worker threads that can schedule asynchronous tasks. Tokio’s own documentation describes the multi-thread scheduler as spawning threads for task scheduling, while also maintaining separate support for blocking work.
Imagine four runtime worker threads:
Tokio Runtime
|
Scheduler
|
┌───────────────┼───────────────┐
│ │ │
Worker 1 Worker 2 Worker 3
│ │ │
Task A Task D Task G
Task B Task E Task H
Task C Task F Task I
Different worker threads can execute tasks at the same time on different CPU cores.
This means Tokio can give us both concurrency and parallelism.
Suppose we have 20,000 asynchronous network tasks.
20,000 Tasks
Many may be waiting.
17,000 waiting for I/O
2,500 waiting for timers
500 runnable
The runtime does not need 20,000 OS threads.
The runnable tasks can be scheduled across a smaller collection of worker threads.
20,000 async tasks
↓
scheduler
┌──────────────┼──────────────┐
↓ ↓ ↓
worker 1 worker 2 worker 3
↓ ↓ ↓
CPU 1 CPU 2 CPU 3
Now return to the juggler analogy from the beginning.
The tasks are the balls.
A Tokio worker thread is closer to one of our jugglers.
One worker can move between many tasks as those tasks become able to make progress.
With several worker threads, several jugglers can work at the same time.
We now have:
many tasks in progress
↓
concurrency
multiple worker threads executing simultaneously
↓
parallelism
The two ideas we separated at the beginning of the article have now come back together.
What if a Tokio task blocks?
There is one final detail we need before we compare these systems directly.
Consider this:
async fn handle_request() {
do_some_very_slow_blocking_work();
do_something_async().await;
}
The function is declared async.
That does not magically make every operation inside it non-blocking.
If do_some_very_slow_blocking_work() blocks the OS thread, then the Tokio worker thread running this task is blocked.
Other tasks assigned to that execution resource may have to wait.
Tokio’s documentation specifically warns that blocking calls or long computations inside a Future can prevent the executor from driving other Futures. Tokio provides spawn_blocking to run blocking operations on threads intended for that kind of work.
Conceptually:
Tokio async workers
|
Async Task
|
blocking operation
|
↓
spawn_blocking
|
↓
Blocking Thread Pool
Once again, we have a hybrid system.
event-driven asynchronous tasks
+
worker threads
+
separate blocking threads
+
operating-system I/O mechanisms
That should sound familiar by now.
Node.js also combines event-driven network I/O with a worker pool.
Tomcat combines network polling with request-processing threads.
Tokio combines asynchronous tasks, an I/O driver, scheduler threads, and threads for blocking work.
The details are very different.
The same fundamental questions keep appearing.
How do we represent work?
How do we represent waiting?
How do we know when waiting work can continue?
Who schedules it?
What happens when something blocks?
How many execution contexts do we really need?
We are now ready to put Node.js, Apache Tomcat, and Tokio beside each other.
Once we do that, the answer to the title of this article starts to become much clearer.
Node.js, Apache Tomcat and Tokio Side by Side
At this point we can finally put the major pieces into one table.
The table is a simplified mental model. Each system has more options and details than can fit here.
| Question | Node.js | Apache Tomcat | Tokio |
|---|---|---|---|
| Common application model | callbacks, Promises, async/await | request-processing threads and Servlet APIs | Futures, tasks, async/await |
| Common representation of waiting application work | suspended continuation associated with a Promise | sleeping or blocked worker thread in traditional blocking code | Future/task in a pending state |
| Network concurrency | event-driven I/O through libuv and OS facilities | NIO polling plus request-processing threads | event-driven I/O through runtime I/O driver and OS facilities |
| Main scheduling mechanism | event loop for JavaScript work | JVM and OS scheduling of worker threads | Tokio task scheduler running tasks on worker threads |
| Can use multiple CPU cores? | yes, through worker threads, processes, native work and other mechanisms | yes, worker threads can execute in parallel | yes, multi-thread runtime workers can execute tasks in parallel |
| What happens when normal application work blocks? | the event-loop thread can stop unrelated JavaScript work from progressing | one worker thread becomes unavailable until the blocking operation finishes | one runtime worker can stop making progress on other tasks assigned to it |
| Common escape for blocking or CPU-heavy work | worker threads or other worker pools | thread pools, dedicated executors, asynchronous Servlet facilities | spawn_blocking, dedicated compute pools, separate threads |
The differences are important.
The common ground is more important for understanding concurrency itself.
All three systems are trying to answer some version of this question:
There are many operations in progress.
Only some of them can use the CPU right now.
What should represent the operations that are waiting,
and how should runnable work get CPU time?
That is the real connection.
Event-Based Concurrency Does Not Remove Threads
We can now clear up another common misunderstanding.
When people first discover event-driven servers, it is easy to imagine that event-based programming replaced threads.
That is not what happened.
Look at Node.js.
event loop
+
libuv worker threads
+
worker_threads
+
operating-system threads underneath
Look at Tomcat.
NIO polling
+
request-processing threads
+
JVM threads
+
operating-system scheduling
Look at Tokio.
async tasks
+
runtime worker threads
+
blocking thread pool
+
operating-system event mechanisms
The event model gives us another way to represent a large amount of waiting work.
Threads are still very useful execution resources.
In many modern systems, events decide when work is able to make progress, while threads provide the CPU context in which that work actually runs.
A simplified event-driven system can therefore look like this:
thousands of operations
↓
some are waiting
↓
events report readiness
↓
runnable work
↓
small collection of threads
↓
CPU cores
That is very different from:
events replaced threads
Events and threads frequently work together.
Thread-Based and Event-Based Concurrency Are Different Ways of Preserving Waiting Work
Let us return to the two mental models we introduced much earlier.
Thread-oriented concurrency:
"Preserve my execution context while I wait."
Event-oriented concurrency:
"Preserve the state I need,
and continue me when I can make progress again."
These two statements explain a surprising amount of what we have seen.
A traditional Tomcat request can preserve the execution context through its worker thread.
The stack remains there.
The blocking call returns.
The next line runs.
Node.js can suspend the logical operation and preserve what is needed for its continuation.
When the Promise settles, JavaScript continues from the appropriate point.
Tokio can represent the suspended computation as Future state.
When progress may be possible, the runtime polls that Future again.
The programmer experiences three different programming models.
The machine still needs answers to the same basic questions.
Where is the state?
What is waiting?
What event makes it runnable?
Who schedules it?
Which thread executes it?
How long can it keep that thread?
These are concurrency questions.
They existed before Node.js.
They existed before Tokio.
They existed before async/await.
The APIs have become better.
The runtimes have become more sophisticated.
The underlying problem remains recognisable.
So Which Model Is Better?
There is no universal answer.
Suppose you have a service with a moderate number of requests and the libraries it depends on are naturally blocking.
A thread pool can be a very sensible choice.
The application code is easy to follow.
The operating system is extremely good at scheduling threads.
Modern machines have large amounts of memory compared with machines from earlier generations.
For many systems, this model works extremely well.
Now suppose you need to maintain tens of thousands of mostly idle network connections.
Most of these connections spend their time waiting for data.
Assigning a full operating system thread to every connection may be unnecessary.
Event-driven I/O can make it possible for a much smaller set of execution threads to manage those connections.
Now suppose the workload is CPU-heavy.
Maybe you are encoding videos.
Maybe you are processing large images.
Maybe you are running complex calculations.
An event loop does not create more CPU capacity.
If the machine has eight cores, only a limited amount of CPU work can physically happen in parallel.
At that point, using multiple execution threads and distributing CPU work across the available cores becomes important.
When deciding how a system should handle concurrency, useful questions include:
- How much of the workload is spent waiting for I/O?
- How many operations may be in progress at once?
- How many of those operations are usually runnable?
- How much CPU work does each operation perform?
- How much memory can be spent on execution contexts?
- Are the libraries we depend on blocking or asynchronous?
- How much complexity are we willing to introduce into the programming model?
- What latency and throughput do we need?
There is no rule saying the entire system must use one model.
As we have already seen, real systems mix them.
Node.js uses events and threads.
Tomcat uses events and threads.
Tokio uses events and threads.
The interesting difference is where each system places the boundary.
Threads Versus Events Is the Wrong Argument
This is where discussions about concurrency can become misleading.
People sometimes talk about these models as if one has to defeat the other.
threads versus event loops
blocking versus async
thread-per-request versus tasks
Real systems are more interesting than that.
A Tomcat server can use event-based socket polling and still execute application requests on worker threads.
Node.js can use an event loop for JavaScript and still send certain work into libuv’s worker pool.
Tokio can multiplex large numbers of asynchronous tasks across several operating system threads and also maintain a separate pool for blocking operations.
The boundaries are different.
The building blocks overlap.
This is why the more useful question is:
How should this system represent work that is waiting, and how should it schedule work that is ready?
For one system, a waiting operation may be a sleeping thread.
For another, it may be a Promise continuation.
For another, it may be a Future returning Poll::Pending.
Eventually, all runnable code has to reach a CPU.
That gives us a more complete picture:
operations in progress
|
┌──────────┴──────────┐
│ │
waiting runnable
│ │
│ ↓
│ scheduler
│ │
│ ↓
│ thread
│ │
│ ↓
│ CPU
│
↓
event / wake-up source
│
└────────────→ runnable again
Different systems put different names on the boxes.
The boxes themselves are much older than the runtimes we are discussing.
What Does Node.js, Apache Tomcat and Rust’s Tokio Have in Common?
We can finally answer the question in the title.
Node.js, Apache Tomcat and Tokio all sit between a large amount of concurrent work and a much smaller amount of actual execution capacity.
They all need to manage work that is currently able to run.
They all need to manage work that is waiting.
They all need some way to decide when waiting work can continue.
They all eventually need threads and CPU time to execute instructions.
Where they differ is in how they represent that work and where they choose to preserve its state.
Tomcat commonly gives a request an operating system thread from a worker pool.
request
↓
worker thread
↓
blocking application code
Node.js commonly represents many concurrent operations through an event loop, callbacks, Promises and continuations.
request
↓
JavaScript
↓
async operation
↓
suspend
↓
event
↓
continue
Tokio represents asynchronous computations as Futures running inside tasks that are scheduled across runtime workers.
task
↓
poll Future
↓
Pending
↓
event
↓
wake
↓
poll again
Under all three, the operating system is still there.
Sockets are still there.
Threads are still there.
Schedulers are still there.
CPU cores are still finite.
Waiting still exists.
The runtimes give us different ways to organise all of this.
The Deeper Lesson
We started this article with a juggler.
One juggler could keep several balls in motion.
That gave us concurrency.
Several jugglers could work at the same time.
That gave us parallelism.
We can now make the analogy a little more useful.
The hard part is not simply keeping many balls in the air.
The system also needs to know which ball needs attention right now.
Some operations need CPU time.
Some are waiting for a network packet.
Some are waiting for a database.
Some are waiting for a timer.
Some are waiting for another task.
Some are ready to continue immediately.
Concurrency is about organising all of these operations so that waiting work does not unnecessarily prevent runnable work from making progress.
Threads give us one way to do that.
Event loops give us another.
Futures and tasks give us another layer of abstraction over the same problem.
async/await gives us a much nicer way to write some of those asynchronous computations.
Thread pools give us a controlled amount of execution capacity.
Operating-system facilities such as epoll, kqueue and IOCP help runtimes discover when I/O operations can make progress.
When you look at the whole picture, the progression becomes easier to understand:
OS Threads
↓
preserve execution contexts
non-blocking I/O
↓
avoid sleeping on one resource
epoll / kqueue / IOCP
↓
efficiently discover I/O progress
event loops
↓
dispatch ready work
callbacks
↓
represent what should happen next
Promises / Futures
↓
represent unfinished computations
async / await
↓
write those computations in a sequential style
Node.js / Tomcat / Tokio
↓
different combinations of these ideas
This is why I do not think the most interesting lesson is:
Node.js uses an event loop.
Tomcat uses threads.
Tokio uses async Rust.
Those statements are useful introductions, but they stop too early.
A more useful understanding is this:
A concurrent system needs a way to represent work, preserve the state of unfinished work, notice when progress becomes possible, and schedule runnable work onto the hardware available to it.
Node.js answers those questions one way.
Apache Tomcat answers them another way.
Tokio answers them another way.
Once you see concurrency at this level, Node.js’s event loop stops looking like a special trick.
Tomcat’s worker threads stop looking like an older idea that event-driven systems somehow replaced.
Tokio’s tasks, Futures, polling and waking stop looking like concepts that appeared specifically for Rust.
They are different parts of a much older conversation about how computers should deal with a simple fact:
A program can have thousands of things in progress while only a handful of them can use the CPU right now.
The job of a concurrency model is to decide what happens to everything else.
