Blog

What is real-time decisioning? Definition, pipeline, and tests

What real-time decisioning actually requires: latency budgets, data prerequisites, and a 6-question test to tell true real-time from fast batch in disguise.

what-is-real-time-decisioning-definition-pipeline-and-tests
Alex Patino Alexander Patino Solutions Content Leader Published September 3, 2026 Read time 31 min read

A customer abandons a cart at 2:14 p.m. and buys the same product from a competitor at 2:31 p.m. But because the first retailer only generates cart-recovery email messages overnight, a reminder and an encouraging coupon don't arrive until the next morning, after the customer has already made the purchase from the competitor. 

At a different company, a card transaction clears at 11:03 p.m. against a fraud score computed from a customer profile last refreshed at 6:00 a.m. The 17 hours of behavior between those two timestamps contained every signal needed to decline the transaction. But the score approved it anyway, because the score was answering a question about a customer who no longer existed.

Both failures came from the same thing. Customer and operational context decays in seconds, and many systems decide on stale data. Batch scoring pipelines, overnight segment refreshes, and scheduled campaign logic all add delay where the situation can change. The dashboard that finds the fraud pattern the next morning is too late to help. This is why real-time decisions are important. 

It doesn’t help to license "real-time" decisioning platforms if the data they use still runs in batch rather than arriving in real time. The engine may evaluate in milliseconds, but that doesn’t help if it’s evaluating segments computed last night, profiles synced hourly, and consent states reconciled weekly. 

So how do you get true real-time decisions? What do they require, and how do you determine whether a system is truly real-time?

What is real-time decision-making?

Real-time decisioning evaluates live context, the current event combined with historical state, to select an action within the window of the interaction, typically in milliseconds. A customer clicks, a transaction posts, a session starts, a sensor fires, and before the transaction completes, the system has weighed who this is, what happened, what is allowed, and what is most valuable, and has acted on the result.

Three components recur across most implementations. 

  1. Signals, or the live event stream, enriched with historical context so the decision sees both what is happening now and everything relevant that came before. A click means something different from a customer with a ten-year history and an open support complaint than from an anonymous first-time visitor, and the enrichment step is what provides that information.

  2. Constraints, such as business rules, eligibility logic, regulatory requirements, and brand policy, affect the decision. Constraints define what the system is permitted to do before any model weighs in on what it should do. 

  3. Selection typically involves multiple models scoring candidate actions concurrently, with an arbitration layer choosing among them based on predicted value, priority, and policy. The output is one action, such as a decision not to act, delivered to the requesting channel before the interaction window closes.

So here’s what real-time decision-making is not:

  • It is not fast batch, because precomputed scores served from a cache with low read latency are still based on stale data, retrieved quickly. 

  • It is not a triggered campaign with a rule tree of a static path through if-then logic, fixed at design time, that fires when an event matches a pattern. 

  • It is not a recommendation widget refreshed nightly. 

The distinction that matters is whether the decision itself is computed inside the interaction, based on live state, with the capacity to choose differently than it would have chosen an hour ago. If the answer is no, the system is automating delivery, not decisions. 

Webinar: Achieving the perfect golden record with graph data for identity resolution

Unlock the power of real-time identity resolution in the AdTech landscape with insights from AWS, Lineate, and Aerospike experts. Discover how graph data models and reference architectures can help you build the perfect golden record for seamless audience targeting.

How a real-time decision pipeline works

The pipeline runs in six stages, with a cost and a failure mode to each one.

Event ingestion and enrichment

An event arrives, such as a page view, a transaction, or an API call, from a channel requesting a decision. The system resolves the identity behind the event and fetches the context that makes the event meaningful, such as the customer profile, recent behavioral history, current session state, eligibility flags, and consent status. 

This stage is dominated by data retrieval, and it uses more of the latency budget in the pipeline. It’s also the most important. If identity resolution fails, every downstream stage relies on a guess; if context retrieval is slow, the budget is used up before anything has been evaluated; if the fetched profile is stale, the decision will be, too. 

Rule evaluation

Constraints run first because they are cheap and they limit the possible decisions. Eligibility, compliance, frequency caps, channel permissions, and suppression rules eliminate candidate actions before any model wastes time scoring them. 

Rule evaluation is fast, typically single-digit milliseconds. It also needs to be maintained; rules accumulate, contradict, and shadow one another over years, and an unaudited rule base eventually causes problems. 

Concurrent model scoring

Remaining candidate actions are scored, usually by multiple models running in parallel such as propensity, value, risk, and churn,each using features that must be fetched or computed in real time. Scoring latency is dominated by feature retrieval, which is a data problem. Its most common problem is feature staleness: a model serving features computed last night produces an outdated score. 

Arbitration and action selection

An arbitration layer reconciles the scores against priorities, expected value, and policy, and selects one action, or none. This is where competing objectives collide: the retention offer and the cross-sell offer and the service message all want the same impression. Arbitration doesn’t use much computation, but requires governance more than performance, because whoever controls the arbitration logic controls the behavior of the system.

Delivery

The selected action returns to the requesting channel: the web page renders the offer, the payment processor receives the approve-or-decline, or the agent desktop displays the prompt. Delivery must complete inside the remaining latency budget. Otherwise, it times out; an action selected too late is no action, and most channels will have already fallen back to a default by the time it arrives.

Outcome capture

What happened, whether it’s accepted, ignored, converted, or charged back, re-enters the system as training signal and as state for the next decision. This stage has no time pressure, which is why many systems don’t take advantage of the opportunity. Without captured outcomes, there is no learning loop, no measurable lift, and no way to know whether the engine is making good decisions or merely fast ones.

In production, the data-retrieval stages of enrichment and feature fetching both take the most time and potentially cause the most problems, while the scoring stages dominate the marketing. 

Real time" means latency budgets

"Real time" is not a speed; it is a deadline. The latency budget means the decision must complete inside the interaction window, and the interaction window varies by channel. 

  • Web personalization has to decide before the page renders, which in practice means keeping the decision call under 100 milliseconds. Anything slower adds directly to the time a shopper waits for content, and the stakes at that scale are measurable. In Milliseconds Make Millions, a study Deloitte Digital ran for Google across 37 brands and more than 30 million mobile sessions, a 0.1-second improvement in mobile site speed was associated with an 8.4% increase in retail conversions and a 9.2% increase in average order value. 

  • Fraud authorization must complete inside the transaction window. 

  • Send-time optimization for email tolerates seconds. 

The budget is set by the channel, not by the vendor, and a system that cannot meet the channel's budget is not real-time for that channel.

Where the milliseconds go

 A 100-millisecond decision budget must cover the network hop from the channel to the decision service, the identity resolution, the profile read, the feature retrieval including multiple reads across multiple stores, the rule evaluation, the model scoring, the arbitration, and the response back across the network. Data retrieval uses most of the budget, and decision logic uses just a little. 

This is why architectures built on cross-system hops, federated queries, or batch-synced data stores have trouble running it in real time: they use up all the time retrieving data. 

The tail is the product

Mean latency of a decisioning system is not a useful metric. The numbers that matter are p99 and p999, or the latency experienced by the slowest 1% and 0.1%of requests, because at scale, latency variability is unavoidable and the tail grows with fan-out. When a decision uses many backend services, the probability that at least one of them is slow rises sharply, so even rare slowness in components slows down the whole system. 

Worse, users don’t experience the tail uniformly. It concentrates at volume peaks such as launches, flash sales, and fraud waves, when decisions carry the most value. A system with an excellent mean and a heavy tail is a system that works except when it matters. 

Webinar: Driving ROI with high performance data infrastructure for AI

Want to get real ROI from your AI investments? Watch the webinar, Driving ROI with high-performance data infrastructure for AI, and see how enterprises use modern data architecture: built on low latency, high-throughput systems to turn AI from pilot projects into real business impact. Watch now and discover how to architect AI data systems that scale, cut costs, and deliver measurable value fast.

What happens when the budget blows

Every real-time system needs a defined answer to the question of what it does when the deadline arrives before the decision. The three options are: 

  • Time out to a default action

  • Serve the last known good decision

  • Skip the decision entirely 

All three are legitimate engineering choices. What is not legitimate is failing to measure how often they happen. An engine that frequently falls back to defaults has degraded to batch behavior, where the default action is made without live context, while every dashboard continues to report the system as healthy and real-time. The fallback rate is the clearest metric of whether a decisioning deployment is real-time.

The data prerequisites 

Most real-time decisioning programs fail because the data prerequisites weren’t met, not because the decision logic was wrong. Four requirements must be met for the project to succeed. Organizations should evaluate those requirements before choosing a platform, not afterward.

Identity resolution rate

Every decision begins by answering “who is this?” If most incoming events cannot be tied to a known profile because identifiers are fragmented across channels, because cookie loss has degraded match rates, or because the identity graph was never built, the engine is guessing. 

Depending on the task, this may not matter. Fraud decisioning extracts value from device and behavioral signals even on anonymous traffic, but next-best-action personalization doesn’t do much without an identity. What percentage of decision-eligible events resolve to a profile, per channel, today? If this isn’t known, the program is not ready for an engine.

Feature freshness SLAs

Every feature a decision uses has an implicit freshness requirement, which varies. An exit-intent intervention needs session signals that are seconds old, but a daily send tolerates features computed overnight. 

A freshness service level agreement commits that a specific feature will reflect reality within a specific window, and your pipelines either meet it or do not. Most organizations have never written down the freshness requirements of their decisioning features, which means they don’t know how well the system meets them. 

Outcome logging

The learning loop requires that every decision's outcome, whether it’s acted on, ignored, converted, or disputed, goes back into the system, attributable to the decision that produced it. Without closed-loop outcome capture, there is no model improvement, no measurable lift, and no answer to the question of whether the program is working. 

Event completeness across channels

A decisioning engine that sees web events but not call-center events, or transactions but not service complaints, isn’t making decisions with a complete view of the customer. The completeness audit asks which behaviors affect the decision, at what latency, and which do not. For example, a cross-sell offer served during an open complaint is almost always an event-completeness failure, not a logic failure.

Implementation veterans consistently report that data integration, building the pipelines, resolving the identities, and meeting the freshness requirements use most of decisioning project timelines, not the decision logic. Instead, audit the four gates, fix the failures, and select the platform whose requirements your data meets. 

Real-time decisions vs. what it replaces

Here are some aspects of the different decision types:

Rules engine

Batch scoring pipeline

Fast batch ("real-time" branded)

True real-time decisions

Latency to act

Milliseconds

Hours to days

Milliseconds to serve

Milliseconds

Decision computed

At design time, by humans

On schedule, by models

On schedule, served fast

At the moment, by models plus rules

Content freshness

Whatever the rules read

As of last batch run

As of last batch run

Live event plus current state

Adaptability

Manual rule changes

Periodic retraining

Periodic recomputation

Closed-loop learning

Typical use

Eligibility, compliance, simple triggers

Propensity lists, segmentation

Cached recommendations, triggered sends

Fraud authorization, next-best-action (NBA), dynamic pricing

Two aspects of this table deserve emphasis. The first is that everything has rules. Every production decisioning system runs rules as guardrails, such as eligibility, compliance, or suppression, around the model layer. The replacement target is not rules but rules-as-the-entire-decision, where every path is human-authored, maintenance scales with complexity, and nothing learns.

The second is the third column. While it’s often branded as “real time,” fast batch is precomputed decisions served quickly: segments scored overnight and cached, recommendations refreshed nightly behind a low-latency API, triggered campaigns whose rule trees were fixed at design time. From the outside, it looks like real-time decisioning, because the serve is fast. 

The difference is that the decision was made based on customer data from hours ago. Fast batch is the most common production state in the market; it is what a large share of "real-time" license spend actually operates, and it is the configuration many organizations are using. 

Are you actually doing real-time decisioning? A diagnostic

 The following tests can be disproved; they work equally well on your own stack and on a vendor's demo, and each one separates decisioning from delivery automation. 

Can it decide within the interaction, under production load?

Measure decision latency, or request-to-action, at production concurrency, not in a demo environment with one user and a warm cache. Use the p99, not the mean. A system that decides in 40 milliseconds at demo load and 400 at peak load is essentially a batch system. 

Can it re-decide within one session?

When new signals arrive mid-session, such as a search, a complaint, or a near-abandonment, does the next decision reflect them, or was the customer's path fixed at session entry? Fixed-at-entry indicates a triggered campaign tree. Re-decision on live signal indicates real-time decisioning.

Is the context live or precomputed?

Trace one decision's inputs. If the deciding features are last night's segment memberships and yesterday's scores, the system is serving fast batch regardless of its serve latency. If the inputs include the current event and state updated within the session, it is deciding on live context.

Does outcome data close the loop automatically, and at what latency?

Ask how an outcome, such as conversion, ignore, or chargeback, reaches the next decision. If the answer involves a quarterly retraining project, the loop is open. If outcomes flow back automatically, ask the loop latency: hours is a learning system; quarters is a reporting system.

Can the system choose silence?

Ask whether doing nothing is a valid outcome, or whether the system is forced to choose a message every time. A system that always sends a message is measuring success by sends, not by outcomes. 

What does it do when the data layer is slow, and do you know your fallback rate?

Ask for the fallback rate, or the percentage of decisions that timed out to a default. If the number exists and is monitored, it is real time. If it isn’t measured, the system's real-time status can’t be proved. 

A system that passes all six is doing real-time decisioning. A system that fails the first three is doing fast batch. 

Read the full Criteo customer story

Powering global, real-time ad bidding with sub-millisecond processing, Criteo scales to 25M transactions per second with Aerospike, reducing carbon footprint by 80%.

The data layer 

The prerequisites section described what your data must be. What matters next is what your data infrastructure must do. An organization can have resolved identities, fresh features, and complete events, and still miss latency deadlines because the substrate serving that data cannot keep up.

Here is what’s required:

  • Reads must complete in sub-millisecond to low-millisecond time under high concurrency, because the profile read and feature fetch sit inside a budget measured in tens of milliseconds and shared with everything else.

  • Writes arrive as an event stream whose peaks are bursty and uncorrelated with read demand, except during traffic spikes. 

A product launch is a write storm and a personalization peak simultaneously. A fraud wave is an event surge and a scoring surge simultaneously. A regional incident is a state-update flood arriving when the suppression logic needs current state. The substrate must handle write peaks without degrading read latency, hold state consistent across regions for customers who interact across them, and exhibit predictable capacity behavior near its limits. 

When the substrate falls short, the result is degraded performance rather than an outage.

  • Feature pipelines lag, so the features the engine reads grow stale. 

  • Read latencies stretch, so decision budgets aren’t met more often. 

  • Blown budgets trigger fallbacks, so the fallback rate climbs. 

  • Fallbacks serve defaults computed without live context, so the engine's decisions converge toward batch behavior. 

And these are some of the results: 

  • A cache layer installed in front of a slower database drifts out of sync with its source, so the engine makes decisions using stale data.

  • Memory pressure forces evictions when traffic peaks fill the cache, so the misses concentrate there. 

  • A hotspot partition means one popular key, such as the trending product or the attacked account, increases tail latency. 

  • During routine infrastructure changes, data reads can become too slow, forcing the system to fall back to default behavior and causing a spike in fallbacks.

Every individual component reports healthy: the database is up, the cache hit rate looks normal, the API is responding, the engine is returning actions, while the system as a whole has stopped being real-time, but without a failure that could alert operators to the problem. 

This is why the database is an important part of real-time decisioning. Remember that most of the time is used by data retrieval. Whether a decisioning program is real-time is based on whether the data layer underneath serves live state, at the tail, under the peaks, every time.

Where real-time decisioning is used

Regardless of the industry, they all fit in a spectrum defined by two aspects: the time available to make a decision and how long that decision remains valid. 

NBA and personalization are at one end. The budget is the page render or the conversation turn, which needs to be under 100 milliseconds for web, a few seconds for an agent desktop, and staleness has an opportunity cost: a stale decision serves an irrelevant offer. 

But fraud and transaction risk are more serious. The decision must complete inside the authorization window, the context must include the transaction in flight, and if it fails, there’s a cost involved. A missed fraud event costs the loss plus the dispute. A false positive costs the sale, the customer relationship, and future revenue, and that adds up. Stale context means either fraud losses or loss of revenue. 

Dynamic pricing sits in between. Budgets range from milliseconds, such as with ad floor pricing, to minutes, such as with ride or delivery pricing. Staleness costs money in both directions: Price too low against current demand and you lose money; price too high and customers don’t buy the service.

Operational decisions such as request routing, resource allocation, and ad serving inside OpenRTB's roughly 100-millisecond bid windows have the tightest budgets in the category. 

Thinking of this in terms of a spectrum makes explicit that "real-time" is a different number in every cell, that the data-freshness requirement scales with the staleness cost, and that an enterprise typically runs several of these engines simultaneously, against the same customers. 

Re-decisioning and learning loops

Two properties separate a decisioning system from a scoring system with good latency:

  1. Re-decisioning, or the capacity to revise the decision repeatedly within a session as new signals arrive, rather than committing to a path at entry. The customer who searched, hesitated, opened a support article, and returned to the cart is four different decisioning contexts; a system that decided once at the beginning is wrong by minute three.

  2. Learning loop: outcomes updating the models continuously, so that the next decision benefits from the last one's result, rather than waiting for a scheduled retraining cycle to incorporate a quarter's worth of evidence.

Loop speed is based on the system, not the model. An adaptive model cannot adapt faster than its outcome data arrives, and outcome data arrives at the speed of the slowest pipeline between the action and the training signal. Observation, decision, and outcome must share a data path. When they run in three systems reconciled by batch jobs, the "continuous" learning loop runs at the cadence of the batch jobs.

What the category never discloses is that the loop has economics, and they are charged up front. 

  • Exploration has a cost: A learning engine serves suboptimal actions to gather the signal that improves future decisions, and every exploratory decision is margin spent on information. The cost is real, quantifiable, and results in an unexplained dip in early performance. 

  • Cold start has a duration: An adaptive model decides poorly before it decides well, and how long that takes depends on traffic volume, outcome frequency, and how rare the events being learned are. 

Neither is an argument against learning systems. Both are line items with known mitigations, such as rules-first ramps that limit the model until it earns autonomy, informative priors, models imported from adjacent domains, and exploration budgets capped as explicit policy. Treating exploration and cold start as managed costs helps a program provide value. 

When the engine breaks: Drift, regime change, and the override path

Every model deployed into a live environment is decaying from the beginning, because the world it was trained on is ending continuously. 

  • The slow version is drift: Customer behavior shifts gradually under the model, accuracy erodes, and the engine's decisions become incrementally worse at a rate too slow to trigger an alarm. 

  • The fast version is regime change: A “black swan” event like a pandemic, a product-line pivot, or a market shock that invalidates the training distribution overnight. For example, when consumer behavior transformed in days during early 2020, machine-learning models running inventory, fraud detection, and marketing systems broke in ways that forced widespread human intervention, because models trained on normal behavior were making decisions for an outdated situation. The problem was not that the models stopped producing decisions, but that they kept producing wrong ones until humans overrode them.

This means the system also needs an override path:

  • Drift detection and alerting, so that decay is observed rather than discovered in quarterly results

  • Confidence thresholds below which the engine defers to a rule or a human. A designed human-override path with audit trails recording who overrode what, when, and why

  • Kill switches scoped per strategy to halt one faulty decision strategy without taking down the engine

  • Rehearsal: To make sure override paths work when you need them

An engine without an override path is unaccountable. And the maturity question for a decisioning organization is not whether its models are good, but whether the organization can detect that its models have stopped being good and act on that detection within hours rather than quarters. 

Deciding to do nothing

Often, it’s assumed that the best action is always an action, but that’s often not true. Choosing silence should be an option, and an engine that cannot produce it is structurally biased toward over-messaging.

The cases where nothing is the right decision are not unusual:

  • Fatigue and frequency governance: The marginal message to an over-contacted customer makes them more likely to unsubscribe or ignore the channel. 

  • Context suppression: Offers during an open complaint, upsells during a service incident, and marketing into a household experiencing a disputed charge are unlikely to be received well.

  • Portfolio restraint: When every candidate action scores below the threshold where acting beats waiting, the best action is to save the customer's attention for when something is worth saying.

Silence is unpopular for a measurement reason, not a strategic one. When the engine sends, and the customer converts, a team gets credit. When the engine holds, and the relationship survives, nothing happens, and nobody gets credit for it. The fix is to track the percentage of decision opportunities resolved as deliberate non-action, and to validate it with holdouts that measure what the suppressed messages would have cost. 

One customer, many engines: Cross-domain arbitration

People often think there’s just one engine deciding for a customer, but that’s often not true. More than likely, it’s several engines, procured by different functions in different years, deciding about the same customer in the same week. 

  • Fraud wants friction on the transaction that marketing wants to convert frictionlessly. 

  • Service wants offers suppressed during an open complaint that the marketing engine cannot see. 

  • Collections and cross-sell pursue the same household simultaneously. 

At the least, organizations should have priority hierarchies such as a fixed ranking, where risk decisions override service decisions override marketing decisions, enforced at the channel. They are crude, and they resolve conflicts by fiat rather than by value, but they are vastly better than nothing.

  • A shared decision layer is the best answer: One arbitration tier through which all domains' candidate actions go, evaluated against a common value framework. It is architecturally clean but hard to do, because the functions that own the engines have to surrender final say to a shared layer, and office politics get involved.

  • Federated engines with a governance contract sit between: Each domain keeps its engine, but all engines honor shared suppression states, a common frequency ledger, and defined precedence rules.

Whatever the architecture, cross-engine frequency governance is needed. Customers can deal with being contacted only so many times, and engines that respect only their own frequency caps overdo this. The frequency ledger must be cross-engine.

Who owns the decision? Governance and operating model

This problem in decisioning programs is organizational. Marketing teams define eligibility metadata, set priorities, and label propositions, and believe they are designing decisions, but their influence ends at data labeling. The arbitration logic that selects among candidate actions is configured by whoever administers the orchestration tooling, which becomes the de facto decision engine of the enterprise while nobody owns decision design as a discipline. The organization has automated its decisions but hasn’t done it on purpose. 

Organizations need to decide:

  • Who owns the arbitration logic, who is authorized to change it, and is that list the people who do it now?

  • How are decision strategies versioned, tested in simulation against historical traffic before they are used in production, promoted through environments, and rolled back when they aren’t operating properly? 

The discipline deserves its name: decisions-as-code, with the same review, testing, and rollback rigor the organization already applies to software. What does the decisioning team actually look like, and who is the decisioning architect, the person who translates between marketing's objectives, data science's models, and engineering's constraints? The role is real, and scarce. 

And what is the cross-functional contract? Marketing owns objectives and propositions, data science owns models and measurement, engineering owns latency and the substrate, and the arbitration layer needs a named owner.

Technology selection is the easy part of decisioning. The organizations that fail mostly fail here, in unowned arbitration logic, unversioned strategies, and a discipline nobody was hired to practice.

Buy, build, or assemble

Given all these aspects to consider, what’s the best path to getting it done successfully?

Buying

Buying a packaged decision hub is fastest. The platform arrives with arbitration, strategy authoring, simulation, and channel connectors built in, and with opinions about how decisioning should work that you will adopt, whether or not they fit. They are powerful, but integration-heavy, expensive to implement, and expensive to keep, with an ongoing cost for integration and specialist staffing. 

A decision hub depends on the latency profile of its data. The hub authors and arbitrates, but still has to deal with the latency budget, which is primarily based on data integration. Buying the hub provides the logic but doesn’t help the data substrate. The buy path makes sense when use-case breadth across marketing, risk, and service justifies the platform's weight and the organization can staff its care.

Building

Building on streaming infrastructure is the path data engineers often use: event streams, stream processing, a feature and state store serving low-latency reads, and your own arbitration layer encoding your own decision logic. But streams and stream processors move and transform state; they do not serve it. A decision is a point read against current state inside a millisecond budget, and no amount of stream throughput substitutes for that read completing at the p99. 

Teams that put together the streaming layer without considering the state store end up with decisions that aren’t real-time, because the component that answers "who is this and what is true about them right now" cannot answer inside the window. The build path has the most control and requires engineering investment, which is justified when latency requirements are extreme, decision logic is proprietary, or the amount and complexity of the data make moving it to a vendor's cloud impractical.

Assemble

Assembling a hybrid is the increasingly common middle: packaged decisioning logic running over your own data substrate, or custom arbitration over managed streaming and storage components. The hybrid means the decision logic and the data layer are separable concerns with different build-versus-buy economics, and that the data layer is the part that determines performance. 

How do you decide? It’s based on the tightest latency budget you must meet at the p99, the breadth of use cases the investment must serve, the engineering capacity you realistically have, where it’s unrealistic to upload your data to a vendor cloud, and how much lock-in you tolerate in a layer this central. The latency requirement eliminates more options than any other feature. 

Five signs you have outgrown Redis

If you deploy Redis for mission-critical applications, you are likely experiencing scalability and performance issues. Not with Aerospike. Check out our white paper to learn how Aerospike can help you.

Decisioning under regulation

Automated decisions about individuals are regulated objects, and you can’t get around this. Under GDPR Article 22, individuals have the right not to be subject to decisions based solely on automated processing that produce legal or similarly significant effects, with narrow exceptions that carry their own safeguard obligations, including the right to obtain human intervention and to contest the decision.

Regulation is also growing. For example, the Court of Justice of the EU now holds that automated credit scoring falls within Article 22's scope.

In U.S. financial services, creditors using complex algorithms or machine learning in credit decisions must still provide adverse-action notices stating the specific, principal reasons for the action, and the regulator has stated that companies are not absolved of these responsibilities because a black-box model made the decision. 

Regulatory issues affect the architecture in three ways:

  1. Explainability becomes a per-decision requirement; the system must be able to reconstruct, for any individual decision, the contributing factors, the confidence, and the rule trace that produced it.

  2. Audit-trail retention becomes a storage and design requirement: decisions, their inputs, and their overrides must be retained and reconstructable.

  3. Consent becomes a real-time data problem of its own; a consent state that is hours stale is a problem if a customer withdrew consent before they were profiled was profiled unlawfully regardless of how fresh every other feature was. 

Another problem is that adaptive models drift by design, while regulators prefer stable, explainable logic. Managing that dichotomy by testing new approaches against the existing one, model versioning, and constrained model classes in regulated decisions needs to be part of the operating model. 

What real-time decisioning won't do

All that stipulated, real-time decisioning isn’t a panacea, and people still need to set the system up. 

It does not replace strategy. Objectives, guardrails, value definitions, and the ethical lines the engine may not cross remain human work; the engine has to work within those limitations, but it can’t set them up itself. 

It does not eliminate rules. Constraints such as compliance, eligibility, and brand policy are permanent infrastructure around whatever the models propose.

It does not fix bad data. Good data becomes fast good decisions, stale data becomes fast stale decisions, and the system doesn’t know the difference.

It runs fraud authorization, pricing, routing, and ad serving as well as martech. 

It does not remove the human. It relocates the human from the operator making individual decisions to architect and oversee designing strategies, monitoring drift, and owning the override path. This is actually harder. 

It does not deliver value just by buying it. Between the contract and the first good decision stand a data-readiness gate, integration work that dominates the project timeline, and a model ramp-up while the engine learns how to do its job. Programs need to plan for that. 

Measuring whether it's working

Once you’ve set up a decisions program, how do you know whether it’s working? There are three categories to measure.

Outcome metrics are the familiar layer: conversion lift, retention, revenue per decision, fraud catch rate weighed against false-positive cost. Necessary, but insufficient alone, because outcome metrics cannot distinguish a working engine from a healthy business.

System metrics cover the database and its infrastructure: the decision latency distribution using p99 rather than the mean, feature freshness against declared SLAs, the fallback rate as the indicator of real-time status, and the do-nothing rate as the health metric for suppression. These metrics detect the degradation of the engine drifting toward batch behavior while outcomes still look fine on trailing averages.

The rigor layer is incrementality. Determining decisioning value requires a control population receiving the old logic, or no decisioning at all, with lift measured against it. Engagement deltas without a holdout measure only that the engine sent things and things happened, not that the engine caused them. 

However, the holdout itself carries a cost and a governance question; it requires serving a control group what the organization believes is worse treatment, which someone must approve, like arbitration logic and override paths. An organization mature enough to run holdouts is an organization mature enough to run decisioning. 

There are also tradeoffs to consider: short-term lift against fatigue cost, exploration expense against exploitation gain, and the fraud catch rate against the false positive rate. Measurement that reports only the numerator of each tradeoff leads to problems. 

Decisioning engines and AI agents

The question every buyer now asks is what happens to decisioning when AI agents plan and act. Is the decisioning engine the tool the agent calls, the governance layer that constrains it, or legacy infrastructure the agent replaces?

The answer is a synthesis with two parts: 

Agents use decisions. An agent executing a multi-step strategy such as resolving a service issue, negotiating a retention offer, or orchestrating a journey needs, at each step, a millisecond-latency evaluation of current context, eligibility, and expected value. Those evaluations are decisions: live context, constraints, and selection, within an interaction window. An agent without a decisioning substrate uses stale data. 

Agents inherit the same substrate requirements, intensified. Agentic latency budgets are decisioning latency budgets with more steps; a five-step agent plan inside a customer interaction is five sequential decision evaluations sharing one interaction window, which tightens the per-decision budget rather than relaxing it. 

The constraint layer also needs modification. The guardrails a decisioning engine already maintains, such as eligibility, compliance, frequency, suppression, the audit trail, and the override path, are what an enterprise needs around agent actions. An agent that acts on customers needs the same things as a model that decides about customers: defined permissions, per-action explainability, a fallback when confidence is low, and a kill switch. 

Aerospike and real-time decisioning 

The decision logic, the arbitration layer, the models, and the governance can all be chosen, changed, and argued about, but underneath every configuration sits a data layer that must answer "who is this and what is true about them right now" in under a millisecond, at the p99, during the write storm, every time. Aerospike is a real-time database engineered to serve live state at sub-millisecond latency with predictable behavior at the p99 tail, and to hold that behavior under adverse conditions: write peaks coinciding with read peaks, datasets in the billions of records, and traffic that spikes when decisions matter. Its architecture removes problems rather than managing them. 

There is no separate cache layer to drift out of sync, because the database itself operates at near cache speed; there is no eviction cascade at peak, and no invalidation lag. Hotspotting and rebalancing are handled by a distribution model designed so routine cluster events don’t add to latency. Strong consistency holds state correctly for multi-region customers. And because it delivers this with a fraction of the server footprint that memory-bound architectures require, the cost of keeping state hot stays linear as the decision volume grows. Predictability under volatility is the design objective the decisioning platform is organized around.

Try Aerospike Cloud

Break through barriers with the lightning-fast, scalable, yet affordable Aerospike distributed NoSQL database. With this fully managed DBaaS, you can go from start to scale in minutes.

Frequently asked questions

Frequently asked questions about real-time decisioning

As fast as the interaction window of the channel. Web personalization and ad decisioning require under 100 milliseconds, fraud must decide inside the transaction authorization window, and send-time optimization tolerates seconds. Measure against the p99 latency, not the average, because the slowest decisions cluster.

A recommendation engine ranks content or products, usually from periodically refreshed models, and answers "what might this person like." A decisioning engine evaluates live context against rules, models, and competing objectives to select one action, or no action, inside the interaction, and answers "what should we do right now." Recommendations are often one candidate input into a decision.

No. A decisioning engine can run on deterministic rules, and rules remain a permanent layer in every deployment. Machine-learning models add adaptability and pattern recognition rules cannot express, but the defining property is deciding on live context within the interaction, not the presence of a model.

Live events from every channel, resolved to persistent customer identities, enriched with historical state and current features, and governed by consent. Requirements are identity resolution rate, feature freshness, outcome capture, and event completeness, and they should be audited before selecting a platform.

RTIM includes decisioning applied to customer interactions across owned channels, usually bundled with journey orchestration. Real-time decisioning is the broader architecture, which also runs fraud authorization, pricing, and operational routing. All RTIM is decisioning; most decisioning is not RTIM.

Test three things. Whether it decides within the interaction at production load based on p99 decision latency, whether it re-decides mid-session when new signals arrive, and whether its inputs are live state rather than precomputed segments. A system that fails these is serving fast batch, regardless of what it calls itself.

In addition to the cost of the software itself, there’s the data work required. The biggest part is integration, such as building event pipelines, identity resolution, and the low-latency data layer, which takes up most of the time, alongside specialist staffing for decision design.

Yes, on streaming infrastructure such as event streams, stream processing, a low-latency state store, and a custom arbitration layer. Engineering organizations do it routinely for fraud, pricing, and routing. Building gives you the most control but requires durable engineering investment; the pragmatic middle assembles packaged decision logic over your own data substrate.

Footnotes

  1. Will Douglas Heaven, "Our weird behavior during the pandemic is messing with AI models," MIT Technology Review (May 11, 2020) https://www.technologyreview.com/2020/05/11/1001563/covid-pandemic-broken-ai-machine-learning-amazon-retail-fraud-humans-in-the-loop/

  2. Regulation (EU) 2016/679 (GDPR), Article 22, "Automated individual decision-making, including profiling" https://gdpr-info.eu/art-22-gdpr/

  3. Bloomberg Law, "GDPR Curbs Use of AI Via Article 22 Decision" (analysis of CJEU Case C-634/21, SCHUFA Holding) https://www.bloomberglaw.com/external/document/X4BBTPFO000000/international-data-privacy-compliance-professional-perspective-g