genread-inference-cost-70s
IN premise — summaries/2026/08/24/mallen-2023-when-not-to-trust-s5-non-parametric-memory-complements.md
Created 2026-08-25T02:58:11+00:00
GenRead (elicitive prompting) incurs approximately 70 seconds per query on GPT-NeoX 20B.
Summary
Using the GenRead elicitive prompting technique on GPT-NeoX 20B adds roughly 70 seconds of processing time per individual query. In practice, this sets a hard latency floor for any workflow built on that pairing, so systems relying on it must be designed around slow turn-around rather than real-time responsiveness, and batch workloads need to budget for that per-query overhead.