Skip to content

Feature vectors for serving

For the complete documentation index see: llms.txt

All documentation pages available in markdown.

In Part 1, the Entity class already defined get_feature_vector(), so you do not need to jump back and edit earlier cells here. This page focuses on how that method works in serving and how to validate its behavior.

Use get_feature_vector

Unlike the other Entity methods that use the Spark connector for batch reads, get_feature_vector uses the Aerospike Python client. It accepts the as_client connection you created on the previous page.

get_feature_vector looks up one record by primary key and calls as_client.select(...) to return only the bins in feature_list.

  1. Run Cell 19 to retrieve one driver’s feature vector.

Cell 19: Retrieve one driver’s feature vector

feature_columns = ["ds_decl_rate", "ds_avg_rating", "da_trips_today"]
features = Entity.get_feature_vector(as_client, "driver", "driver_042", feature_columns)
print(f"Features for driver_042: {features}")
Expected output
Features for driver_042: {'ds_decl_rate': 0.061, 'ds_avg_rating': 4.72, 'da_trips_today': 11}
  • as_client: Aerospike Python client connection
  • etype: entity type (for example, "driver")
  • eid: entity ID (for example, "driver_042")
  • feature_list: bins to retrieve (for example, ["ds_decl_rate", "ds_avg_rating", "da_trips_today"])
  • Returns {feature_name: value}, or None if the entity doesn’t exist

Test with known drivers

Try a few drivers to see the range of feature values. Recall from Part 2 that drivers with ds_decl_rate >= 0.10 were labeled as higher decline risk.

  1. Run Cell 20 to retrieve feature vectors for test drivers.

Cell 20: Retrieve feature vectors for test drivers

feature_columns = ["ds_decl_rate", "ds_avg_rating", "da_trips_today"]
test_drivers = ["driver_005", "driver_042", "driver_087"]
for driver_id in test_drivers:
fv = Entity.get_feature_vector(as_client, "driver", driver_id, feature_columns)
if fv is None:
print(f"{driver_id}: not found")
continue
risk = "higher" if fv["ds_decl_rate"] >= 0.10 else "typical"
print(f"{driver_id}: decline_rate={fv['ds_decl_rate']:.3f}, "
f"rating={fv['ds_avg_rating']:.2f}, "
f"trips_today={fv['da_trips_today']} ({risk} risk)")
Expected output
driver_005: decline_rate=0.048, rating=4.81, trips_today=3 (typical risk)
driver_042: decline_rate=0.061, rating=4.72, trips_today=11 (typical risk)
driver_087: decline_rate=0.167, rating=4.22, trips_today=9 (higher risk)

Your values will differ, but you should see a mix of typical and higher-risk drivers based on the threshold.

Measure retrieval time

Run 100 individual lookups (one per driver) to get a stable latency measurement. Each lookup retrieves the same 3-feature vector for one driver.

  1. Run Cell 21 to benchmark 100 feature vector retrievals.

Cell 21: Benchmark 100 feature vector retrievals

import time
feature_columns = ["ds_decl_rate", "ds_avg_rating", "da_trips_today"]
timings = []
for i in range(1, 101):
driver_id = f"driver_{i:03d}"
start = time.perf_counter()
Entity.get_feature_vector(as_client, "driver", driver_id, feature_columns)
elapsed = (time.perf_counter() - start) * 1000
timings.append(elapsed)
timings.sort()
print("Feature retrieval benchmark: 100 lookups (3 features each)")
print(f" p50: {timings[49]:.2f} ms")
print(f" p95: {timings[94]:.2f} ms")
print(f" p99: {timings[98]:.2f} ms")
Expected output
Feature retrieval benchmark: 100 lookups (3 features each)
p50: 0.22 ms
p95: 0.38 ms
p99: 0.45 ms

Retrieval stays sub-millisecond across the sample. Each call is a single key lookup, and Aerospike returns the requested bins directly.

You now have a fast feature retrieval method. Next, you’ll connect it to the trained model to make predictions.

Feedback

Was this page helpful?

What type of feedback are you giving?

What would you like us to know?

+Capture screenshot

Can we reach out to you?