Announcing pi-xai-ws v1.0.0

Published:
Keywords: ai

Contents

Introduction

I’m happy to announce pi-xai-ws↗ v1.0.0, the first stable release of a Pi↗ package I’ve been using to talk to Grok. It intercepts Pi’s built-in xAI Responses models and sends their turns over xAI’s official Responses WebSocket instead of HTTP.

What pi-xai-ws Does

Pi already knows how to call Grok. This package changes how those calls travel to xAI:

  • One WebSocket is reused for a Pi session, and model calls are serialized on that connection.
  • By default, every call still sends Pi’s complete local history with store: false.
  • If you opt into stored-response mode, later calls send only the newest items. A durable checkpoint lets a reconnect or a new Pi process resume.

It reuses the SuperGrok OAuth credentials already stored in Pi. Pi 0.84 or newer is required.

Cache Coherency

xAI prompt caching is prefix-based. For a long Grok agent thread to stay cheap and fast, later turns need to keep hitting that cached prefix. Compaction, rewritten history, or a rejected stored reference can still force a cold replay. In ordinary use, cache reuse has been very consistent.

I measured my own Pi sessions with stored continuation enabled. Across 21 sessions that is 2,787 Grok Responses calls, about 672 million prompt tokens:

  • 96.1% of those prompt tokens came from cache.
  • The median call was 99.5% cached.
  • 92.9% of calls were at least 90% cached.
  • Only 0.9% of calls were fully uncached, which is expected for a first turn or a cold replay.

What keeps the prefix stable is one socket per session, cache affinity from the session ID, and stored continuation so later turns send only new items. A durable checkpoint survives idle close, token refresh, and Pi restart. Older screenshots become placeholders once the request exceeds an 8MB image budget, and they stay omitted so they don’t reappear in the wire prefix.

Stored continuation is off by default because xAI retains saved Responses state for 30 days. Enable it only if that is acceptable. With storage off, each call sends the full local history, and cache reuse on SuperGrok is weaker.

Error Handling and Reliability

Cache reuse matters less if a long job stops halfway through. The transport tries to keep going when the socket dies, xAI is overloaded, or Grok gets stuck.

  • Quiet or expired connections are retried, including before xAI’s 25-minute socket limit.
  • If Grok stops after thinking only, or starts repeating itself, the run continues instead of ending there.
  • If a stored response is too large to keep, already-streamed output is saved and later calls fall back to full local history.

For longer Grok jobs, I use agent-level retries with maxRetries: 5 and baseDelayMs: 3000, leaving provider retries at 0. Those Pi settings are in the README.

Getting Started

Install from npm:

pi install npm:@mwolson-org/pi-xai-ws

Or try it for one run:

pi -e npm:@mwolson-org/pi-xai-ws

To opt into stored-response continuation, create ~/.pi/agent/pi-xai-ws.json:

{
  "storeResponses": true
}
  • Frontend
  • Mobile
  • Backend