Dedicated servers for streaming and transcoding
Video streaming infrastructure has two distinct jobs: preparing video files for playback and delivering those files to viewers. Transcoding converts an input such as a 4K H.264 upload, RTMP live feed, or camera stream into multiple resolutions and bitrates. Delivery serves HLS playlists, MPEG-DASH manifests, MP4 files, thumbnails, and segments to viewers over HTTP.
A dedicated server is valuable because video workloads consume resources continuously. A single poorly optimized 4K-to-1080p software transcode can use several CPU cores for its entire duration. A busy live channel may generate new HLS segments every two to six seconds, while hundreds of viewers create sustained outbound traffic. Bare-metal servers provide reserved CPU time, direct NVMe performance, predictable network capacity, and no noisy-neighbor contention from other virtual machines.
Common streaming workloads
- Video-on-demand: user uploads are encoded into 1080p, 720p, 480p, and mobile-friendly renditions for HLS or DASH playback.
- Live streaming: RTMP, SRT, or WebRTC input is packaged and transcoded into adaptive HLS streams.
- Game, event, and education broadcasts: recurring live channels need stable ingest, monitoring, recording, and replay storage.
- IP camera systems: NVR software can record RTSP feeds, generate low-resolution previews, and retain footage for a defined period.
- Media platforms: self-hosted video portals, training libraries, paywalled content, and internal communications platforms need controlled storage and access rules.
VPS versus dedicated server
A VPS is sufficient for a small video library, a single low-traffic live channel, a development environment, or lightweight remuxing. Start with a VPS if the platform performs mostly direct file delivery, stores fewer than 500 GB of media, has predictable low traffic, and runs no more than one or two occasional 1080p software encodes. For example, a 4 vCPU, 16 GB RAM, and 240 GB NVMe VPS can run Nginx, a small HLS catalog, and background FFmpeg jobs if encoding is rate-limited.
Move to dedicated hardware when transcoding is sustained, viewers are numerous, live latency matters, or the platform handles business-critical content. Dedicated servers are especially appropriate for more than three concurrent 1080p software transcodes, 4K encoding, multi-camera recording, more than 5 TB of monthly egress, or workloads that cannot tolerate CPU-steal time. They are also preferable for mail notifications, databases, monitoring, and streaming services that must coexist without competing for limited virtual CPU resources.
GPU encoding can substantially increase stream density, but it changes output quality and operational design. NVIDIA NVENC, Intel Quick Sync Video, and AMD hardware encoders are efficient for live ABR ladders and high-volume delivery. CPU encoding with x264 or x265 generally offers better compression efficiency at a given bitrate, particularly for archived video-on-demand content. Test representative source files before committing to a codec or hardware encoder.
Capacity planning: transcodes, storage, and bandwidth
Do not size a streaming server by viewer count alone. CPU and GPU requirements are driven primarily by simultaneous transcode jobs, source resolution, codec, frame rate, target ladder, and encoding preset. Network requirements are driven by viewers and selected bitrates. A 1080p HLS viewer watching a 5 Mbps rendition uses about 2.25 GB per hour before protocol overhead. One hundred viewers at 5 Mbps require approximately 500 Mbps of sustained outbound capacity.
Storage calculations should include source files, encoded renditions, HLS segments, thumbnails, logs, and a working area for FFmpeg. A one-hour 1080p source at 8 Mbps is about 3.6 GB. An adaptive bitrate package with 1080p, 720p, and 480p renditions may use a similar or greater amount of total storage depending on bitrates and codecs. Keep at least 20% of NVMe capacity free so temporary files, filesystem metadata, and database writes do not slow down.
For up to two simultaneous 1080p software transcodes, a 4 vCPU / 16 GB / 240 GB NVMe VPS can work; at three to eight jobs, use a dedicated 8-core / 32 GB / 480 GB NVMe server, and use 16 dedicated cores or a GPU-equipped server for larger queues.
| Simultaneous 1080p software transcode jobs | vCPU or CPU cores | RAM | Disk | Monthly bandwidth |
|---|---|---|---|---|
| 1–2 jobs; small HLS catalog | 4 vCPU | 16 GB | 240 GB NVMe | 5 TB |
| 3–8 jobs; several live channels or VOD queue | 8 dedicated CPU cores | 32 GB | 480 GB NVMe | 10 TB |
| 9–20 jobs; high-volume VOD or multi-channel live platform | 16 dedicated CPU cores | 64 GB | 1 TB NVMe | 20 TB |
| 20+ jobs; GPU-assisted encoding cluster | 16 cores plus hardware encoder GPU | 128 GB | 2 TB NVMe plus object storage | 50 TB+ |
These figures assume H.264 1080p inputs and reasonably balanced FFmpeg presets. HEVC, AV1, 4K, 60 fps, denoising, subtitles, and multiple output renditions can increase CPU time sharply. For GPU-assisted pipelines, validate the GPU encoder session limits, available VRAM, codec support, and measured quality at your target bitrate.
Recommended server architecture
Operating system and base services
Use a current long-term-support Linux distribution such as Debian 12 or Ubuntu 24.04 LTS. Keep the streaming stack separated into services: Nginx handles HTTP delivery, FFmpeg handles encoding, a queue system controls jobs, PostgreSQL or MariaDB stores metadata, and monitoring watches service health. Docker Compose is suitable for a small installation; systemd services or Kubernetes may fit larger teams with established operational experience.
Place public HTTP delivery behind Nginx. Limit direct access to upload endpoints, use TLS certificates, and require signed URLs or authenticated sessions for private content. For live ingest, expose only the required RTMP, SRT, or WebRTC ports and restrict publisher credentials. Do not make administrative dashboards or storage mounts publicly accessible.
HLS and DASH packaging
HLS is broadly compatible with browsers, smart TVs, and mobile devices. Use short segments for lower startup and recovery time, but avoid making segments unnecessarily short because segment creation adds filesystem and HTTP overhead. For most standard HLS deployments, six-second segments and a playlist containing six to ten segments are practical starting values. Low-latency HLS requires a different tuning approach, including partial segments, compatible players, more HTTP requests, and careful origin performance testing.
For video-on-demand, encode a bitrate ladder instead of serving one oversized file. A practical H.264 ladder might include 1080p at 4.5–6 Mbps, 720p at 2.5–3.5 Mbps, and 480p at 1–1.5 Mbps. Actual values depend on content: sports, games, and high-motion scenes require more bitrate than slides, lectures, or talking-head video.
Step-by-step deployment recommendations
1. Prepare the dedicated server
Create a non-root administrator, install updates, and enable a firewall before deploying applications. Use SSH keys, disable password authentication after testing key access, and permit only ports needed for SSH, HTTPS, and selected ingest protocols.
sudo apt update && sudo apt -y upgrade && sudo apt install -y nginx ffmpeg ufw2. Verify storage and mount paths
Use NVMe storage for active uploads, HLS segments, transcode scratch space, and databases. Create separate directories for source uploads, temporary work, published media, and backups. Do not store the only copy of customer uploads on the same filesystem as temporary encoding files.
sudo mkdir -p /srv/media/{uploads,work,hls,vod} && sudo chown -R www-data:www-data /srv/media3. Generate an adaptive bitrate HLS package
Test a representative source file before automating jobs. The command below creates 1080p, 720p, and 480p H.264 renditions with AAC audio and an HLS master playlist. Use a job queue so the server does not start more encodes than its CPU budget supports.
ffmpeg -i input.mp4 -filter_complex "[0:v]split=3[v1][v2][v3];[v1]scale=-2:1080[v1o];[v2]scale=-2:720[v2o];[v3]scale=-2:480[v3o]" -map "[v1o]" -map 0:a -c:v:0 libx264 -b:v:0 5500k -maxrate:v:0 6000k -bufsize:v:0 8250k -c:a aac -b:a 128k -map "[v2o]" -map 0:a -c:v:1 libx264 -b:v:1 3000k -maxrate:v:1 3300k -bufsize:v:1 4500k -c:a aac -b:a 128k -map "[v3o]" -map 0:a -c:v:2 libx264 -b:v:2 1200k -maxrate:v:2 1320k -bufsize:v:2 1800k -c:a aac -b:a 96k -f hls -hls_time 6 -hls_playlist_type vod -master_pl_name master.m3u8 -var_stream_map "v:0,a:0 v:1,a:1 v:2,a:2" /srv/media/vod/stream_%v.m3u84. Configure Nginx for efficient delivery
Set correct MIME types, allow byte-range requests for MP4 delivery, and set cache headers appropriate to the content. Immutable completed segments can be cached longer than live playlists. Keep live playlist caching short so players receive current segment references.
sudo nginx -t && sudo systemctl reload nginx5. Add HTTPS and access controls
Use TLS for player pages, APIs, and media endpoints. For private libraries, issue short-lived signed URLs from your application instead of exposing predictable media paths. Rate-limit login and upload endpoints, but avoid overly aggressive limits on segment delivery because a normal player requests many small files.
sudo certbot --nginx -d stream.example.com6. Monitor the actual bottleneck
Watch CPU load, per-process memory, disk latency, filesystem free space, network throughput, HTTP errors, and FFmpeg exit codes. A load average near the number of physical CPU cores during active encoding is not automatically a problem, but growing encode queues, high I/O wait, dropped live frames, or player buffering are signs that capacity is insufficient.
sudo apt install -y sysstat bmon && iostat -xz 1Performance optimization
- Use a queue with concurrency limits: reserve one or two CPU cores for Nginx, the database, and the operating system. An 8-core server should not normally run eight demanding x264 jobs at once.
- Choose sensible FFmpeg presets: slower x264 presets improve compression but consume more CPU. For time-sensitive live video, start with
-preset veryfastor-preset faster; evaluate quality with real footage. - Use CRF for archived quality targets: CRF encoding is useful for source preservation or VOD preparation, while capped bitrate or CBR-like settings are more predictable for live delivery.
- Keep active segments on NVMe: frequent small HLS writes are a poor match for slow network filesystems or overloaded SATA disks.
- Separate origin and delivery at scale: once egress becomes the largest cost or bottleneck, keep the origin server focused on ingest and transcoding while using a CDN or additional cache nodes for public delivery.
- Use hardware encoding carefully: NVENC or Quick Sync can encode more concurrent streams per watt, but measure output quality and ensure the selected GPU supports the required H.264, HEVC, AV1, 10-bit, or HDR features.
- Optimize keyframe alignment: align GOP size with segment duration. At 30 fps with six-second segments, start with a 180-frame GOP so HLS segment boundaries are easier to place at keyframes.
Common pitfalls to avoid
- Buying for storage but ignoring egress: a server may have enough disk for thousands of videos yet run out of monthly transfer after a popular stream.
- Running unlimited FFmpeg processes: this causes CPU saturation, memory pressure, high disk I/O wait, delayed segment generation, and player buffering.
- Using one bitrate for every viewer: high-bitrate-only video excludes mobile and weak-network users; adaptive bitrate ladders reduce buffering.
- Keeping uploads, scratch files, and backups together: a full disk can stop active transcodes and corrupt operational assumptions. Use quotas, retention policies, and off-server backups.
- Serving unprotected media paths: private content needs authentication, authorization, and expiring URLs or tokens.
- Skipping monitoring: test alerts for disk capacity, failed jobs, high packet loss, high CPU temperature where available, and failed TLS renewal.
- Assuming every codec plays everywhere: test target browsers, smart TVs, mobile apps, and embedded players before deploying HEVC or AV1 as the only option.
Choosing a Valebyte server
Valebyte VPS plans are a sensible starting point for a small HLS site, test pipeline, or a single low-volume channel. Choose a Valebyte dedicated server for sustained CPU encoding, high transfer requirements, sensitive media libraries, multi-channel live delivery, or a production system where predictable resources matter. Start from measured queue length, storage growth, and outbound Mbps rather than selecting hardware solely by the size of the original video files.