Executive Summary
EventStream by Markets IO (MIO) is a service-oriented, real-time data distribution platform. Providers and publishers group their content into services, and each service offers channels: requestable, manageable streams of related events, each identified by a topic, that entitled users subscribe to.
In Access Point Gateway (APG) mode, EventStream takes services from connected providers and publishers over TCP and delivers them to connected client applications. An APG can act as a simple router between providers and users, but it can also add value along the way, with optional pipeline engines for last-value caching, load balancing across providers and quality-of-service functions such as conflation and delay. APGs can also be cascaded, for example with a caching APG tier on-boarding providers and a second APG tier delivering to users in a DMZ, giving a scalable, resilient and secure layered design.
Internally, each APG pipeline has a gateway engine at the top, optional value-add engines in the middle and an access point engine at the bottom. Each of these tiers can be scaled out across threads: Gateway threads for input and output scaling, Channel threads for caching and channel scaling, and Access Point threads for user fan-out scaling. This allows EventStream to be shaped to both the workload and the hardware it runs on.
In September 2026, CJC carried out performance testing of EventStream 1.20.0 in APG mode on bare-metal Supermicro servers powered by the AMD EPYC 9575F processor, comparing a series of thread configurations under sustained, high-volume load.
Objectives
The objective of this document is to provide a detailed analysis of EventStream APG performance across different thread configurations, comparing placement within a single NUMA node against placement spanning multiple NUMA nodes. This report includes:
- Specifications: the bare-metal server, CPU, network and operating system platform.
- Configurations: BIOS, operating system, network adapter and CPU binding applied to optimise performance.
- Testing Methodology:
- Throughput: maximum sustainable update rates.
- CPU Utilisation: efficiency of Gateway, Channel and Access Point threads.
- Network Bandwidth: inbound and outbound traffic handling.
- Latency: end-to-end data delivery performance.
- Thread Model Comparison: the effect of thread counts and NUMA placement on throughput and latency.
Summary of Key Findings
- Latency: Headline Result – The Single Node Throughput configuration (six Gateway threads, no Channel threads and three Access Point threads on a single NUMA node) delivered the best latency of the testing. At 100,000 updates per second it achieved a mean latency of 61 µs, a median of 60 µs and a 99th percentile of 84 µs, with a standard deviation of just 9.4 µs. The 99th percentile stayed below 100 µs up to 200,000 updates per second. At 540,000 updates per second the median was still 78 µs and the 99th percentile 159 µs, with 30% of updates delivered in under 50 µs and the fastest in 25 µs. At 2.4 million updates per second, 78% of its maximum tested rate, latency remained consistent, with a 99th percentile of 476 µs and a maximum of 513 µs. Even when driven to 3.06 million updates per second, with its Gateway threads near saturation, no sample exceeded 1 ms. All figures are full end-to-end measurements, including two traversals of the switched network.
- Throughput – EventStream delivered 8 million updates per second outbound from a single NUMA node (400,000 inbound, 20 clients) and 16.2 million updates per second from a multi-node configuration (810,000 inbound, 20 clients), saturating one and two 10GbE links respectively. A dedicated ingest test sustained 3.06 million updates per second inbound and outbound.
- Network as the Constraint – In all fan-out tests the 10GbE network saturated before EventStream reached its CPU limits. During the best-performing multi-node test the APG server remained 83% idle.
- Latency Under Full Fan-Out Load – With the client-facing network links saturated, the Single Node OOB configuration achieved a mean latency of 141 µs with a 99th percentile of 245 µs and no samples above 1 ms. At 810,000 inbound and 16.2 million outbound updates per second, the best multi-node configuration achieved a mean of 261 µs.
- Thread Model – At high load, moving from the three-tier model to the two-tier model by removing Channel threads roughly halved mean latency (591 µs to 313 µs). With caching now in the Gateway tier, doubling Gateway threads from three to six then relieved Gateway CPU pressure (around 90% to around 53%), reducing mean latency further to 261 µs and maximum latency from 1,593 µs to 1,064 µs.
- Platform – The high-clock AMD EPYC 9575F, combined with a low-latency operating system profile and NUMA-aware thread binding, provided a strong and predictable foundation for EventStream.
About CJC:
CJC is the leading market data technology consultancy and service provider for global financial markets. CJC provides multi-award-winning consultancy, managed services, cloud solutions, observability, and professional commercial management services for mission-critical market data systems. CJC is ISO 27001 certified, enabling CJC’s partners the freedom to focus on their core business.
Connect with us on social media: LinkedIn | Twitter
|
What's Inside?
- Chapter 1 - About CJC.
- 1.1 - What CJC Does.
- 1.2 - CJC Solutions Include.
- Chapter 2 - Introduction.
- Chapter 3 - Executive Summary.
- 3.1 - Objectives.
- 3.2 - Key Findings.
- Chapter 4 - Technical Information.
- 4.1 - Hardware Platform.
- 4.2 - Server Overview.
- 4.3 - Test Environment and Topology.
- 4.4 - BIOS Configuration.
- 4.5 - CPU and NUMA Layout.
- 4.6 - Network Adapter Configuration.
- 4.7 - Operating System Configuration.
- 4.8 - Software.
- 4.9 - EventStream APG Threading Model.
- Chapter 5 - Testing Overview.
- 5.1 - Configurations Tested.
- Chapter 6 - Preparations for Performance Tests.
- Chapter 7 - Summary of Results.
- 7.1 - Latency Across Update Rates.
- Chapter 8 - Throughput Testing Results.
- 8.1 - Common Information.
- 8.2 - Single Node OOB.
- 8.3 - Multi Node 1.
- 8.4 - Multi Node 2.
- 8.5 - Multi Node 3.
- 8.6 - Single Node Throughput.
- Chapter 9 - Latency Testing Overview.
- 9.1 - Latency Results Summary.
- 9.2 - Latency Tests.
- Chapter 10 - Conclusion.
- Chapter 11 - Notice Regarding Obligations and Conditions.
|