How data moves between two machines - for real production systems

May 21st, 2026 - 11 min read

Abstract data packets streaming between two glowing machine nodes

Copying a file feels trivial until the file is 40GB, the link is shared with a thousand clients, one side is a phone on 4G, and "done" still has to mean correct.

This is how I think about data transfer between two machines - the same model I use for APIs, backups, ETL, object sync, and media pipelines in production.

Libraries change. Physics does not. If you can name the bottleneck (latency, bandwidth, backpressure, disk, TLS handshakes) and the success contract at the app layer, the tool choice becomes secondary - and reviews get sharper.

What "transfer" actually means

On paper: bytes leave machine A and appear on machine B.

In reality you are coordinating five jobs at once:

  1. Encode - turn your thing into bytes (JSON, Protobuf, Parquet, raw file)
  2. Move - push those bytes across a path that lies, drops, reorders, and delays packets
  3. Control - decide how fast to send so nobody melts (flow control / backpressure)
  4. Verify - prove B got what A meant (checksums, lengths, schemas)
  5. Survive - retry, resume, and stay correct when the world is half broken

Miss any one and you get silent corruption, timeouts that look like "network bugs," or a dashboard that says 100% while the consumer is drowning.

Bandwidth, latency, and throughput are not synonyms

People say "the network is slow." That sentence hides three different failures.

WordPlain meaningProduction question
LatencyTime for one signal to go A → BHow long until the first useful byte?
BandwidthHow fat the pipe is (bits/sec)What's the theoretical ceiling?
ThroughputHow much useful work you actually getAfter protocol tax, retries, and waits?
GoodputApplication-level useful data rateAfter encryption, headers, and retransmit?

A path can have great bandwidth and terrible latency (cross-region fiber). Or low bandwidth and fine latency (a local serial link). Design for the bottleneck you actually have.

Rule of thumb: small messages are latency-bound. Bulk transfers are bandwidth-bound - until you introduce head-of-line blocking, tiny buffers, or chatty request/response patterns that turn bulk into thousands of round trips.

Bandwidth-delay product (rough intuition):
  in-flight bytes ≈ bandwidth × RTT

If your window is smaller than that, you will never fill a fat, long pipe -
no matter how "fast" the NIC datasheet looks.

The stack under every "send()"

You almost never talk "machine to machine" raw. You talk through layers:

App bytes
  → serialization / framing
  → TLS (usually)
  → TCP or QUIC / UDP
  → IP routing
  → NIC → cable / radio → NIC
  → reverse the stack on the other side

Each layer adds headers, buffering, and failure modes. Production debugging is often: which layer is lying?

  • App says "sent" - OS accepted into a socket buffer, not "B has it"
  • TCP says "acked" - B's kernel got bytes, not that your app processed them
  • HTTP 200 - the request finished; your business invariant might still be wrong

Always define the success contract at the application layer.

"We got a 200" is not "the ledger is correct." If the consumer crashed after writing 90% of a file, or retried and appended twice, your HTTP status lied politely. Success belongs to checksums, idempotency keys, and committed application state - not to status codes alone.

TCP is a reliability machine with a speed governor

TCP gives you a byte stream that is ordered and (usually) complete. It does that with sequence numbers, ACKs, retransmits, and congestion control.

What engineers must feel in their bones:

  • Windowing - the sender may only have N unacked bytes in flight. Slow ACKs or small windows throttle you even on a fat pipe.
  • Congestion control - packet loss is treated as "the path is full." Burst loss on Wi‑Fi can tank throughput even when average capacity is fine.
  • Head-of-line blocking - one lost packet stalls later bytes in that stream. HTTP/2 over TCP feels this; HTTP/3 / QUIC reduces it with independent streams.

UDP is not "faster TCP." UDP is "you own reliability." Fine for video frames and game state. Dangerous for ledgers unless you rebuild what TCP already solved.

Backpressure: the difference between a pipeline and a bomb

Load is not just "how many MB/s." Load is who is slower.

If A produces faster than B can consume:

  • buffers fill
  • memory climbs
  • latency spikes
  • then you drop, block, or crash

Healthy systems make pressure visible:

  • TCP windows shrink (kernel backpressure)
  • HTTP/2 / gRPC stream windows fill
  • queues grow and workers scale - or shed load on purpose
  • producers block or apply token buckets instead of unbounded Promise.all

Unbounded concurrency is how a "simple sync job" takes down Postgres and the network card on the same afternoon.

// Anti-pattern: unbounded fan-out
await Promise.all(urls.map((u) => fetch(u)))

// Healthier: bounded concurrency (sketch)
async function mapPool<T, R>(
	items: T[],
	limit: number,
	fn: (item: T) => Promise<R>,
): Promise<R[]> {
	const results: R[] = []
	let i = 0
	async function worker() {
		while (i < items.length) {
			const idx = i++
			results[idx] = await fn(items[idx])
		}
	}
	await Promise.all(Array.from({ length: limit }, () => worker()))
	return results
}

Serialization and framing are part of the transfer

Bytes are not "the data." Bytes are a contract.

Ask:

  • Fixed schema (Protobuf, Avro) or flexible (JSON)?
  • Length-prefixed frames or newline-delimited records?
  • Endianness and versioning when language runtimes differ?
  • Is the unit a file, a row batch, an event, or a blob?

In production I prefer:

  • Explicit framing - know where a message starts and ends
  • Versioned schemas - evolve without breaking old consumers
  • Checksums at message or chunk boundaries - catch truncation early

JSON over HTTP is fine until it isn't: CPU for parse, no partial resume, and easy-to-miss size limits at every proxy.

Compression trades CPU for wire time

Gzip, zstd, brotli, and columnar formats (Parquet) exist because wire and disk are often more expensive than cycles.

Compress when:

  • payload is repetitive text or structured logs
  • cross-region egress costs matter
  • disk or object storage volume hurts

Skip or lighten compression when:

  • data is already compressed (JPEG, MP4, encrypted blobs)
  • CPU is the scarce resource (tiny instances, huge fan-out)
  • latency of tiny messages matters more than bytes

Measure. Blind gzip on already-compressed media wastes time and heat.

How production systems actually move bulk data

Different jobs, different tools - same physics.

1. Request/response APIs

Good for small, authoritative reads/writes. Bad as a bulk pipe if you page one row at a time across continents.

Prefer: batch endpoints, cursors, ETags / If-None-Match, pagination that cannot skip or duplicate under concurrent writes.

2. Streams and chunked uploads

Multipart upload (S3-style), tus, resumable gRPC streams, chunked HTTP bodies.

You need:

  • chunk size tuned to RTT and buffer memory
  • resume tokens after disconnect
  • idempotent commit of the final object

3. File and disk oriented tools

rsync, scp, rclone, database dump pipes, snapshot replication.

They shine when the unit of truth is a filesystem tree or a backup artifact. Watch for: sparse files, permissions, partial writes, and "success" that only means the process exited zero.

# Integrity after a bulk copy - size alone is not enough
sha256sum big.dump
# Compare on the other side before you delete the source

4. Message buses and logs

Kafka, NATS, SQS, Pulsar - transfer over time with fan-out.

Here "between two machines" becomes "between many consumers with different speeds." Retention, offsets, and consumer lag are your transfer metrics.

5. Object storage as the meeting point

Often the best "transfer" is: A uploads to S3/GCS/R2, B downloads (or processes in place). You trade direct A↔B coupling for durability, retries, and scale.

Reliability patterns that survive production

Exactly once is a marketing phrase

Across two machines and a network, you get at-least-once or at-most-once unless you design idempotency keys, dedupe stores, or transactional outboxes.

Design for at-least-once + idempotent consumers. That is the adult version.

Checksums end arguments

MD5/SHA for integrity, CRC32C for fast chunk checks. Compare size and hash. Never trust only "bytes written == bytes read" across flaky middleboxes.

Timeouts need a story

Connect timeout ≠ read timeout ≠ overall deadline. A stuck transfer with an infinite read timeout holds sockets, file locks, and human patience hostage.

Partial failure is the default

A finished mid-file. B wrote 90%. The job retried and appended again. Now you have corruption that looks like "encoding bugs."

Prefer write-to-temp + atomic rename, or multipart commit, or versioned keys.

Load: what changes when many transfers share a path

One 1 Gbps flow is not ten 1 Gbps flows. Reality:

  • NICs and interrupts contend
  • disk on B becomes the bottleneck before the network does
  • TLS handshakes dominate short transfers
  • NAT / conntrack tables fill under many short connections
  • fairness: one fat elephant flow can starve latency-sensitive APIs on the same host

Production habits:

  • separate bulk and interactive traffic where you can (queues, QoS, dedicated endpoints)
  • cap concurrent transfers per tenant
  • prefer fewer long-lived connections over reconnect storms
  • watch p99 latency of the interactive path while bulk jobs run

If your backup window kills checkout, you do not have a networking problem. You have a scheduling problem.

Security is not optional wrapping paper

Anything that crosses a trust boundary needs:

  • TLS (or equivalent) in transit
  • authn/z for who may pull or push
  • encryption at rest if the meeting point is shared storage
  • least privilege credentials with rotation
  • audit of who transferred what, when

A "temporary open port for scp" becomes a permanent incident.

A practical checklist before you ship a transfer path

  1. What is the unit of transfer (byte, message, file, snapshot)?
  2. What is the success signal at the app layer?
  3. Latency-bound or bandwidth-bound?
  4. Who is slower under peak load - and how does backpressure show up?
  5. Can we resume after a mid-transfer death?
  6. Are retries idempotent?
  7. How do we detect silent corruption?
  8. What happens to interactive traffic while bulk runs?
  9. What's the blast radius if credentials leak?
  10. Which metrics prove it: throughput, lag, error rate, retry rate, p99?

If you cannot answer those, you do not understand the transfer yet - you only have a script that works on a happy day.

The mental model I keep

Data transfer is a control problem disguised as a copy problem.

Two machines, one unreliable path, asymmetric speed, and a correctness bar that does not care that the demo worked on localhost.

When I design or review a system, I do not start with the library name. I start with load shape, failure shape, and the definition of done. The bytes follow.

That is the difference between "we sent the file" and a production data path you can trust when traffic spikes, a region blips, and someone still expects the numbers to reconcile in the morning.

Doctor discovery product graphic

Doctor Finder & Instant Booking

Help patients find specialists near them and book into real hospital systems - so your marketplace captures demand instead of losing it to call centers.

See the business story
Secure video streaming graphic

Secure Video Hosting at Scale

Private streaming that feels first-party - with YouTube-backed storage and a lean middle tier designed for high concurrency and low cost.

See the business story
Hamidul Islam
Written by Hamidul Islam

Hamidul Islam is a product engineer focused on performance, systems thinking, and the path from hardware into software. He builds product systems that stay fast under pressure and shares what he learns here.

Learn more about Hamidul

Have a question about this article?

Send me a note via the contact page or schedule a call.

Contact me

If you found this article helpful.

You will love these ones as well.