Prompt prediction API guide

Predict the next user prompt or autocomplete a draft with Clone API. Request examples, prediction modes, JavaScript and Python clients, context and response validation.

Clone API provides prompt prediction through POST https://api.clone.is/v1/predictions. It predicts the next user prompt from conversation context or completes a draft already being written. Call it from your backend with a runtime app key, then show the suggestion in your product for the user to review.

Prompt prediction and prompt autocomplete

ModeDraftBehaviorExample use
next_promptEmpty draft.textSuggest the next user message from the current contextPropose a follow-up after an assistant response
complete_draftNonempty draft.textSuggest a continuation of the user's current draftAutocomplete in a chat composer or custom editor

Both modes can use recent messages, known artifact text and product-owned preferences. Optional Clone personalization requires an explicit user connection. Basic predictions work without a Clone end-user account.

Prompt prediction proposes the user's input; it does not run your assistant or complete the task itself. Keep your existing model, chat backend and send controls. For React components or a headless editor controller, use Clone SDK.

Make a next-prompt prediction

Create an app in the free sandbox. Store its runtime key as CLONE_APP_KEY on your backend. Run this server-side request, replacing the example user ID with the subject from your authenticated session:

Shell
curl --fail-with-body --max-time 15 https://api.clone.is/v1/predictions \
  -H "Authorization: Bearer $CLONE_APP_KEY" \
  -H "Content-Type: application/json" \
  --data-binary '{
    "request_id": "next-prompt-example-001",
    "user_id": "your-authenticated-user",
    "session_id": "thread-001",
    "mode": "next_prompt",
    "draft": { "text": "", "revision": 0 },
    "context_revision": "1",
    "messages": [
      { "role": "user", "content": "Draft a launch announcement." },
      { "role": "assistant", "content": "Here is an introduction and invitation to try it." }
    ]
  }'

A suggested response contains completion text. An abstained response has an empty completion and zero prediction units. Predictions may fail or abstain; keep normal input and sending usable. The HTTP quickstart shows illustrative response bodies and the identity fields to validate.

Complete a draft instead

Make a new request with a fresh request ID, mode: "complete_draft" and the user's nonempty draft. These are the fields that change in the full request above:

JSON
{
  "request_id": "draft-example-002",
  "mode": "complete_draft",
  "draft": { "text": "Make the announcement ", "revision": 1 }
}

This fragment is not a complete request. Preserve the authenticated subject, session and current context from the full request. Advance the draft revision on edits and the context revision when the conversation, artifact or preferences change. Do not display or accept results for an older state.

Choose HTTP, JavaScript or Python

The API works with any server that can make HTTPS requests. For a JavaScript/TypeScript backend, install @clone-ai/prompt-prediction@0.7.1 and use CloneClient from its /server export. For Python 3.11+, install clone-sdk==0.2.1 and use CloneClient or AsyncCloneClient from clone_sdk. Reuse one client for the application lifetime and close it at shutdown. The SDK quickstart shows both server clients.

Use an authenticated product proxy between a browser or mobile app and Clone. Derive user_id on the server and apply your product's rate and usage limits. Keep runtime app keys out of client bundles.

What must I validate before displaying a prediction?

Check the echoed request, session, draft, context and connection identities and the expires_at time. Expiry is Unix time in seconds. Hide abstentions, expired results and errors. Cancel abandoned requests where possible, and clear candidates after an account or connection change.

Use a new request ID for new input. To recover an unknown outcome after a timeout, reuse the same request ID and identical body rather than automatically creating another billable request. See the API reference and reliability guide for exact limits, error codes and cancellation behavior.

How much does prompt prediction cost?

The free sandbox includes 1,000 suggestions once per developer account, with no card or automatic paid upgrade. Production predictions use usage billing; SDK installation is free. Consult current API and SDK pricing for rates, budgets and production setup.

Continue with adding prompt suggestions to your product, the HTTP quickstart or the OpenAPI contract.