The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Foresea Forecasting listing page.
Foresea (foresea.ink) is an autonomous prediction market intelligence platform and empirical research framework. It combines real-time probability forecasting across Polymarket and Kalshi, statistical edge discovery, multi-model shadow trading tournaments, an autonomous 19-tool ReAct execution agent, and a public Model Context Protocol (MCP) server.
The repository also serves as the artifact for research studying how explicit rationale instructions, evidence injection, and reasoning structures affect LLM forecasting behavior and calibration on Metaculus-style questions.
Deployed on Google Cloud Run with high-concurrency scaling, streaming responses, and continuous delivery:
/): Real-time landing view with market radar, model-vs-market gap highlights, and interactive walkthroughs./chat, /chat/:id): Conversational interface with streaming rationale, news evidence citation, and multi-turn market analysis./edge, /edge/markets): Ranked live Foresea-vs-market pricing discrepancies across Polymarket and Kalshi, backed by calibration and lead-time scores./edge/mtm): Continuous mark-to-market valuation and PnL tracking for resolved and open market predictions./edge/agentic): Multi-model autonomous trading tournament. Independent $10,000 shadow accounts for each model (Gemma 4 26B, Qwen 3.8 27B, GPT-OSS 120B, GLM 5.3, GLM 5.3 Flash, DeepSeek V4 Flash, MiniMax M3, Llama 3.3 70B) executing real-time paper trades, tracking equity curves, and logging hourly cycle health./trade): Non-custodial, client-side order preview and execution terminal with Cloud KMS envelope encryption for Polymarket and Kalshi credentials./watchlist): Follow specific markets with automated daily digest emails./forecast/:share_id): Shareable forecast permalinks with rationale cards and provenance.When attach_evidence is true and no news_articles are supplied, /predict
fetches and ranks current news evidence from GDELT, Google News RSS, and Stooq by
default, injects it into the model prompt, and returns the selected
evidence_articles with the forecast. Supplying news_articles skips automatic
retrieval and uses the caller-provided evidence.
The response includes both the forecast and the evidence used by the model:
Use evidence_sources when a client only needs the source list and links. Use
evidence_articles when a client needs the article-level details that were
attached to the model prompt. rationale and model_rationale are generated by
gpt-oss-120b and explain why the model chose its answer and confidence.
When market_probability is supplied, market_analysis is computed
deterministically from the model probability and the market-implied probability.
Foresea continuously evaluates state-of-the-art LLMs against real-world prediction markets, tracking statistical edge, calibration accuracy, and shadow portfolio performance across three dedicated views:
/edge or /edge/markets)buy_yes, buy_no, hold).static/track_record_live.json, calculating Brier scores, Expected Calibration Error (ECE), and lead-time skill once markets resolve./edge/mtm)/edge/agentic)gemma-4-26b-a4b-it)qwen3-8-27b)gpt-oss-120b)glm-5-3) & GLM 5.3 Flash (glm-5-3-flash)deepseek-v4-flash)minimax-m3)llama-3.3-70b-instruct)crowd-follow (no-LLM consensus control)The local crypto micro-market model in src/analyzing_llm_rationale/crypto_5m.py
is built for 5-minute UP/DOWN markets where the goal is profitable selective
trading, not constant action. It combines:
Each forecast returns predicted_outcome, probability_up,
component_probabilities, model-vs-market edge, and a fee-aware strategy.
The strategy only recommends a trade when net expected value clears fees and the
configured no-trade threshold.
Use fold_aggregate and evidence_quality before risking capital. If selection
is unstable or holdout PnL is weak, the correct profitable action is to abstain.
--benchmark-log appends a compact JSONL record for tracking whether the
selected threshold and model mode keep working across benchmark runs.
Resolve completed markets against Binance candles:
The resolver returns pending before expiry and resolved afterward with
actual_outcome, resolved_price, and prediction_correct.
Record and resolve paper signals over time:
The signal log is the running dataset for model improvement: each record stores
the forecast, recommendation, later actual_outcome, correctness, and
pnl_per_contract for actual buy_up/buy_down paper trades. Use
--signal-summary to audit whether resolved paper trades are positive after
fees; trade_ready stays false until the configured trade count, PnL, and hit
rate thresholds are met. Use --dry-run with --paper-loop to preview signals
without writing the log.
Production is served from the custom domain:
The Cloud Run service name, project ID, and region are set at deploy time via gcloud run deploy.
Required runtime environment:
SCADS_AI_API_KEY: Secret Manager secret used by hosted model calls.MODEL_DEVICE=cpu: production Cloud Run runs the CPU image.CUSTOM_DOMAIN=foresea.ink: redirects *.run.app requests to the public domain.GOOGLE_CLIENT_ID: Google OAuth web client ID used by /auth/config.GITHUB_CLIENT_ID / GITHUB_CLIENT_SECRET: GitHub OAuth app credentials. The
OAuth app's callback URL must be the site origin (e.g. https://foresea.ink/).
When unset, the "Continue with GitHub" button is hidden and /auth/github
returns 503. Sign-in also works with Google and email/password.SESSION_SECRET: long random string used to sign browser session JWTs and
derive domain-separated, non-reversible references for authenticated analytics.
Rotating it starts a new attribution cohort; it never exposes account emails.The OAuth client must allow these JavaScript origins:
To update non-secret environment variables without replacing the existing
SESSION_SECRET, use --update-env-vars:
Verify the deployed auth config and health endpoint:
The server is built to scale horizontally on Cloud Run:
/auth/register, /auth/login). Passwords are stored as salted
PBKDF2-HMAC-SHA256 hashes; accounts live in Cloud Datastore.REDIS_URL is set, so they are
shared across instances; otherwise they fall back to per-instance in-memory
state and fail open. /predict (non-personalised requests), evidence
retrieval, and /extract URL fetches are cached; public GETs send
Cache-Control.| Var | Default | Description |
|---|---|---|
REDIS_URL | unset | Memorystore/Redis URL. Shares cache + rate limits across instances. |
PREDICT_CACHE_TTL | 600 | Cache TTL (s) for non-personalised /predict responses. 0 disables. |
EVIDENCE_CACHE_TTL | 900 | Cache TTL (s) for evidence retrieval. |
EXTRACT_CACHE_TTL | 3600 | Cache TTL (s) for /extract URL fetches. |
LOCAL_CACHE_MAX | 1024 | Max entries in the in-memory fallback cache. |
SEARXNG_URL / TAVILY_API_KEY / SERPER_API_KEY / BRAVE_API_KEY | unset | Enable web search as an evidence source. A self-hosted SearXNG is preferred when set, then Tavily, Serper, Brave. Tavily/Serper have free no-card tiers. When none is set, evidence comes from GDELT, Google News, and RSS. |
NEWSAPI_KEY | unset | Enables NewsAPI as an evidence source. |
GET /track-record serves the public forecast track record. The heavy tick loop
does not run on Cloud Run: .github/workflows/track-record-tick.yml runs hourly
on GitHub Actions, updates data/track_record_store.json as the source-of-truth
entity store, writes the public aggregate to static/track_record_live.json, and
commits both files back to main. At runtime, Cloud Run fetches the committed
aggregate from raw GitHub, falling back to the bundled file and then the static
backtest in static/track_record.json.
The Action discovers short-to-medium-horizon Polymarket/Kalshi markets in
separate close-date bands (2-7, 7-14, 14-30, 30-60 days by default) and
calls /predict once per newly snapshotted market/model. If /predict is
protected, set the GitHub secret PREDICT_API_KEY; no server-side
/track-record/tick endpoint is required. TRACK_RECORD_TOKEN is optional and
only enables the agent-enrolled market bridge.
The default scheduled forecast job is deliberately cost-capped: it runs every 6
hours, snapshots at most 2 markets per venue, and forecasts only
gpt-oss-120b plus the no-LLM crowd-follow baseline. Use the manual workflow
dispatch input reforecast_each_tick=1 for a one-off full refresh instead of
forcing every scheduled run to reforecast all open markets.
The homepage market desk uses GET /radar, which is derived from
static/track_record_live.json and its edge_board. Radar highlights current
model-vs-market gaps and keeps the first screen fast by reusing the committed
track-record aggregate instead of scanning venues on every page load.
Raise the Cloud Run throughput ceiling (no idle cost while min-instances=0):
For the lowest-cost public deployment, keep the service on request-only CPU, scale to zero, and cap burst scale-out. This is the profile used by the deploy workflow. Startup CPU boost stays enabled because it reduces cold-start latency without keeping an idle instance warm:
Measure deployed forecast latency after each runtime change:
If cold starts still dominate, raise --min-instances to 1 as an explicit
latency/cost tradeoff.
Market search runs in-process in the main API. The optional Go marketd
microservice is build/test-only in GitHub Actions and is not deployed to Cloud
Run by default.
CI pushes commit-tagged Docker images to Artifact Registry on every deploy. Keep
the docker repository cleanup policy active so old images do not accumulate:
The policy deletes images older than 7 days, keeps the newest 5 versions per
package, and always keeps the main tag.
Docker builds run in GitHub Actions, not Cloud Build; no Cloud Build trigger or staging bucket is required for the normal deploy path.
Once max-instances > 1, provision Memorystore for Redis (billable) and set
REDIS_URL so rate limiting and caching stay correct across instances:
See additional Kalshi and Polymarket endpoints for historical data, account pagination, order management and native exchange streams.
The public Cloud Run API is the easiest integration target. It accepts forecasting questions and returns a typed forecast, model rationale, and optional evidence articles. It is built for resolvable forecasts, not general Q&A.
GET /: landing desk, real-time market radar, and interactive workflow demo.GET /chat, GET /chat/{id}: conversational forecast studio with streaming rationales and evidence.GET /edge, GET /edge/markets: live market edge board with statistical gap rankings and Kelly sizing.GET /edge/mtm: mark-to-market performance of open positions across prediction venues.GET /edge/agentic: multi-model autonomous trading tournament, equity curves, cycle health, and paper trade tape.GET /trade: non-custodial trading terminal for Polymarket and Kalshi with envelope encryption.GET /watchlist: tracked favorite markets with daily digest notifications.GET /forecast/{share_id}: public read-only forecast share permalink.GET /embed/forecast/{share_id}: lightweight iframe-embeddable forecast widget.GET /widget.js: drop-in web component <foresea-card> for publishing live forecasts.GET /health: service health check and operational status.POST /predict: public probability prediction endpoint with optional evidence retrieval.POST /agent/analyze: orchestrated end-to-end analysis of a live market with custom skills and ReAct loops.GET /agent/scan: venue scanner identifying mispriced markets ranked by statistical edge.GET /radar: homepage market desk payload derived from the live track record.GET /track-record: public live track record and historical calibration statistics.GET /track-record/digest: shareable markdown summary of the live track record.GET /pr-agent: opt-in agent-to-agent outreach packet for Foresea discovery.GET /markets/polymarket: fetch live normalized Polymarket quotes, orderbooks, and liquidity data.GET /markets/kalshi: fetch live normalized Kalshi quotes, strike ranges, and ticker metadata.GET /mcp/: public remote Model Context Protocol (Streamable-HTTP) endpoint.GET /.well-known/mcp/server.json: public MCP discovery manifest.POST /analytics/visit: privacy-preserving page visit tracking (linked only to non-reversible references).POST /analytics/event: funnel event recording (forecast_completed, watchlist_add, share_created, digest_sent).GET /analytics/events/summary: 30-day aggregate product analytics summary.POST /forecasts/share: generate an explicit public forecast share ID.GET|POST /chat/conversations: cloud conversation sync for authenticated users.GET|POST|DELETE /favorites: watchlist management and tracking.GET|PUT|DELETE /trading/connections/{platform}: KMS-encrypted per-user exchange connection credentials.POST /trading/preview: dry-run order normalization and limit collar verification.POST /trading/orders: live order submission with explicit two-step user confirmation.GET /trading/portfolio: authenticated balances, open positions, resting orders, and execution fills.POST /trading/orders/{audit_order_id}/reconcile: venue order status and fill reconciliation.DELETE /trading/orders/{audit_order_id}: cancel resting orders at venue.Anonymous chats stay in browser localStorage. Signed-in users sync
conversations through /chat/conversations, while watchlist tracking uses
FavoriteMarket entities exposed through /favorites and /favorites/prices.
The favorites digest runs from .github/workflows/favorites-digest.yml via
scripts/favorites_digest.py.
Forecast sharing is opt-in: clients call POST /forecasts/share to create a
public GET /forecast/{share_id} page. Do not expose full private chat history
in shared forecast views.
POST /agent/analyze runs the whole pipeline autonomously: resolve the market
(fetch a live Polymarket/Kalshi price when an identifier is given) → gather
evidence + forecast → price the edge → run any custom skills →
recommend. It returns one structured report.
Custom skills are your own analysis steps — each runs as an extra model pass
over the question, forecast, and evidence, and comes back as a named section in
the report. Provide a question directly, or a platform + market identifier
(slug/market_id for Polymarket, ticker for Kalshi). Pass history (prior
turns) for multi-turn follow-ups — with history, short follow-ups like "why?" or
"what about June?" are answered in context. BYOK fields (openrouter_api_key,
openrouter_model, provider_base_url) apply here too.
The report includes recommendation (buy_yes/buy_no/hold/no_market_price),
edge, model_probability, market_probability, thesis, evidence_sources,
and pipeline (the ordered steps that ran).
Foresea agents utilize an autonomous ReAct (Reason + Act) loop with dynamic plan formation, tool selection, reflection, and JSON error recovery:
Forecasting & Statistical Edge:
forecast: Calibrated probability forecasting with confidence and rationales.get_market: Normalized quote and market metadata lookup across Polymarket and Kalshi.scan_markets: Discover live markets filtered and ranked by model-vs-market edge.batch_quotes: High-throughput multi-venue quote aggregation.search_evidence & web_search: Live multi-source retrieval (GDELT, Google News, SearXNG, Tavily, Stooq).track_record & edge_board: Historical calibration metrics and open ranked alpha opportunities.market_leaderboard: Track record rankings of top prediction market traders.Venue Orderbooks & Microstructure:
exchange_status: Kalshi exchange status, market trading state, and operational schedule.orderbook: Live bids and asks orderbook depth for Kalshi tickers or Polymarket tokens.market_tags: Polymarket category tags and market classification taxonomy.price_history: Historical price points, timeseries, and OHLC candlesticks.live_data: Real-time sports statistics, play-by-play data, and live game feeds.polymarket_meta: Event series hierarchy, resolution rules, and community discussions.recent_trades: Real-time trade tape and prints (prices, contract sizes, timestamps).Execution & State Management:
place_trade: Immediate-or-Cancel (IOC) paper execution against live orderbook quotes with shadow balance updates.manage_notes: Scratchpad state persistence across ReAct reasoning turns.fetch_api: Safe, sandboxed HTTPS retrieval for external data verification.Automatic Calling Aliases: TOOL_ALIASES in agent_capabilities.py normalizes alternative LLM naming conventions (e.g., http_get → fetch_api, candlesticks → price_history, comments → polymarket_meta, trades → recent_trades, leaderboard → market_leaderboard).
Every signed-in call to POST /agent/analyze (including the streamed endpoint)
also creates a private AgentRun. It retains a bounded, secret-free input
snapshot, lifecycle timeline, model report, and any review-only trade handoff.
Use GET /agent/runs for the newest operator timeline and
GET /agent/runs/{run_id} for one full report. The snapshot intentionally
excludes provider keys, browser credentials, conversation history, and raw
custom-skill instructions. An Agent Run is research only: even when it has a
trade handoff, it cannot create, size, or submit an order; the user must still
create and explicitly confirm a durable Trade Run in the terminal.
Signed-in users can copy a public Foresea model from the Agentic board. The copy
is saved under the user's account as an immutable version-1 research recipe;
it contains the public source model and analysis instruction only—never the
source agent's private context, shadow-account history, exchange connection,
order size, or trading permission. Use POST /agent-profiles/copy with an
allowlisted source_agent_id, then pass the returned agent_profile_id to
POST /agent/analyze.
When a profile is selected, the server resolves the profile's model and
instruction itself, ignores client BYOK/provider/model overrides, and forces
the fixed research pipeline (no tool loop or trade tool). The resulting report
returns its profile ID, source, version, and research_only mode for
reproducibility. A profile may prepare the existing review-only trade handoff,
but it cannot create or submit an exchange order; a signed-in user must still
create a durable Trade Run and explicitly confirm PLACE REAL ORDER in the
trading terminal.
GET /agent/scan lists live markets on a venue, forecasts each, and returns the
ones whose model-vs-market gap clears min_edge, ranked by |edge|.
Params: platform (polymarket or kalshi), limit (markets to analyse, max 8),
min_edge (default 0.1), evidence_top_k. Each market runs a full forecast, so
it's bounded by limit and the result is cached briefly. Response: {platform, scanned, opportunities: [{question, market_url, market_probability, model_probability, edge, recommendation}]}. In the web app, the desk's
"⚡ Scan Polymarket for mispriced markets" button calls this.
Foresea exposes a public remote MCP server at:
It is advertised for discovery at:
The remote MCP server is a thin tool layer over the public API. It exposes:
foresea_forecast: produce calibrated probability forecasts with rationale and news evidence.foresea_analyze_market: evaluate a specific Polymarket/Kalshi market with model-vs-market edge & thesis.foresea_scan_markets: scan live markets ranked by model-vs-market disagreement.foresea_batch_quotes: fetch multi-market quotes across venues in one roundtrip.foresea_check_run: check background execution status for long-running market research runs.foresea_edge_board: top open trading opportunities ranked by statistical edge.foresea_track_record: public accuracy, Brier score, ECE, and calibration metrics.foresea_debate_market: conduct adversarial multi-agent debate (Bull vs. Bear vs. Risk Judge).foresea_optimize_portfolio: calculate optimal Fractional Kelly capital allocations across open edges.foresea_feed_latest: real-time alpha feed combining live market edges, agent trades, and leaderboard rankings.foresea_exchange_status: inspect Kalshi exchange status (trading active flag) and operating schedule.foresea_orderbook: fetch live bids and asks orderbook depth for Kalshi tickers or Polymarket tokens.foresea_market_tags: fetch active category taxonomy and tags from Polymarket.foresea_price_history: fetch historical price points or OHLC candlesticks.foresea_live_data: fetch real-time sports game statistics, play-by-play data, and live event feeds.foresea_polymarket_meta: fetch event series listings, community discussion comments, or sports metadata.foresea_recent_trades: fetch recent executed trade tape / prints (prices, sizes, timestamps).foresea_market_leaderboard: fetch top profitable trader leaderboard and volume rankings.foresea://edge-board, foresea://markets/trending, foresea://track-record, foresea://openapi.json.foresea_market_risk_prompt, foresea_calibrate_hypothesis, foresea_forecast_prompt, foresea_system_prompt.See docs/mcp_commercial_guide.md for full harness setup guides (Claude Code, Google Antigravity, OpenAI Codex, Cursor, Windsurf, OpenHands, Smithery.ai).
Foresea provides ready-to-run client integrations across popular developer and trading surfaces:
scripts/foresea_telegram_bot.py): Interactive bot supporting /forecast <q>, /edge, /analyze <ticker>, /track, and automated subscriber edge alerts.
scripts/foresea_discord_bot.py): Posts rich Discord embeds to announcement channels on schedule.
<foresea-card>)Embed live interactive prediction market forecasts into any blog, news site, or Substack with a single script tag:
An opt-in automated execution runner (scripts/live_trader_bridge.py) connecting Foresea's statistical edge signals to live prediction venues (Polymarket & Kalshi) with strict risk management guards:
GET /pr-agent?audience=mcp returns an opt-in outreach packet that other agents,
MCP catalogs, and tool directories can quote when introducing Foresea. It includes
the one-liner, install command, MCP/OpenAPI links, talking points, and an explicit
no-spam policy.
For operator-run cold outreach to explicit agent endpoints, prepare a target list
and use the local runner. It dry-runs by default and only sends with --send:
Target file shape:
The public API returns the outreach packet; it does not expose an unauthenticated
message-sending relay. The scheduled GitHub Action
.github/workflows/pr-agent-outreach.yml runs every 5 minutes against
data/pr_outreach_targets.json, sends with --send, and records contacted
targets in data/pr_outreach_state.json so repeated scheduled runs do not
re-contact the same agent. For a literal always-running local process, run:
Header values can reference GitHub Actions secrets via environment variables, for
example "Authorization": "$PR_AGENT_TARGET_AUTH".
Seeded automated targets:
https://agentndx.ai/api/submit) — public MCP/A2A/x402 review form.https://mcp.directory/api/submit-server) — public JSON submit route.https://mcpub.dev/mcp) — public MCP JSON-RPC submit tool.Additional listing work that is not suitable for the scheduled HTTP sender lives
in data/pr_manual_targets.json. Current manual/GitHub target: mcp.so issue
https://github.com/daodao97/chatmcp/issues/213.
uvx (Claude Desktop / Antigravity / Codex)For OpenClaw, also add this to the target agent's workspace guidance:
A runnable end-to-end demo (scan → forecast → edge) is in
examples/foresea_agent_demo.py.
Use https://foresea.ink/mcp/ directly in MCP clients that support remote
Streamable HTTP servers. For clients that still require a local stdio command,
run the wrapper locally.
The repo targets Python 3.10+ because the official MCP Python SDK requires it.
To create a repo-local Python 3.11 MCP environment with uv:
That lightweight install avoids pulling the full inference dependency stack
(notably Torch/CUDA) when all you need is the MCP wrapper. In a full development
environment, pip install -e ".[mcp]" is also valid.
MCP client config example:
For a local HTTP MCP endpoint:
Connect MCP clients to http://127.0.0.1:8787/mcp. If a private deployment
requires auth, set FORESEA_API_KEY or pass --api-key; the wrapper forwards it
as X-API-Key.
Quick verification:
Pull the current market-implied probability straight from a venue, then feed it
into /predict as market_probability to compute an edge.
Both return a normalised quote:
probability is null for unpriced/illiquid markets. Quotes are cached briefly
(MARKET_CACHE_TTL, default 30s).
Foresea can submit guarded prediction-market orders, but live execution is
disabled by default. Keep this separate from /agent/analyze: the agent can
recommend buy_yes/buy_no, but order submission requires a signed-in user,
an encrypted exchange connection, FORESEA_ENABLE_BYO_TRADING=true,
execute=true, and the exact confirmation phrase PLACE REAL ORDER.
The browser sends connection credentials only to PUT /trading/connections/{platform}.
Foresea validates them, generates a unique data-encryption key for that one
user/venue connection, and encrypts the credential payload locally. Cloud KMS
wraps the data key using authenticated user/venue context; Datastore receives only
the ciphertext, wrapped data key, and KMS key metadata. The KMS root key never
enters the service process. Foresea never returns credentials to the browser and
rejects inline venue_credentials on preview and order requests.
Create a dedicated KMS symmetric ENCRYPT_DECRYPT CryptoKey and give only the
Cloud Run service account roles/cloudkms.cryptoKeyEncrypterDecrypter on that
key. Configure its fully qualified resource name, not a secret value:
Cloud KMS key rotation is transparent to existing wrapped data keys. The service uses the primary key version for a new connection and KMS selects the needed older version when decrypting an existing one.
Install the optional SDKs in production with:
The Docker image installs trading, so Cloud Run only needs secrets/env vars.
If version-1 connection records already exist, deploy the KMS configuration and
keep the old FORESEA_CREDENTIALS_ENCRYPTION_KEY Secret Manager value available
only during migration. Existing records migrate lazily on their first authenticated
use, or migrate the full set from an environment with Application Default
Credentials and Datastore access:
The command reports counts only and never outputs credentials. Once no version-1
records remain, remove FORESEA_CREDENTIALS_ENCRYPTION_KEY from Cloud Run and
Secret Manager.
Check encrypted account connection metadata (no secrets are returned):
Connect one account over TLS (the payload is encrypted before persistence):
Preview a Kalshi order without execution:
Submit a live order only after reviewing the preview:
For Polymarket, pass the CLOB token_id for the exact outcome, or pass
slug/market_id plus outcome and Foresea will resolve the token id from the
public market record. Limit orders use quantity as shares. Market-buy orders
use max_cost as USD spend when supplied and remain blocked unless
FORESEA_ALLOW_MARKET_ORDERS=true.
After submission, use the audit ID returned by /trading/orders to reconcile
the current venue state instead of assuming a submission was filled. The trade
terminal also exposes this flow, including an explicit CANCEL OPEN ORDER
confirmation before it cancels a remaining resting order.
New terminal submissions use a durable /trading/runs record: Foresea saves a
validated order plan, requires a second exact confirmation to execute that saved
plan, and atomically claims it before contacting a venue. This prevents duplicate
orders from concurrent tabs or Cloud Run instances. Run state follows the linked
audit order when a fill, cancellation, or rejection is reconciled.
Every live submission now passes a second server-side preflight immediately before the venue call. It fails closed when Foresea cannot obtain a fresh market quote and a current portfolio snapshot, or when any of these limits would be crossed:
GET/PUT /trading/guardrails: users may set stricter
limits or pause all new live orders, but cannot increase the platform caps.FORESEA_TRADING_KILL_SWITCH=true: blocks every new live submission without
touching reconciliation or cancellations.The trailing-day budget is deliberately worst-case notional newly risked,
not a misleading synthetic P&L figure. Filled positions are measured from the
venue portfolio snapshot before a new order; exact realized daily P&L remains a
separate accounting/reporting concern. Guardrail passes, blocks, policy changes,
and reconciled fill/rejection/cancellation transitions are appended to
GET /trading/guardrails/events without credentials or order payloads. Configure
the existing SMTP_* and ALERT_* settings to receive operator emails for
submission-unknown, rejection, fill, and platform-kill-switch events.
Production ceilings are environment variables; conservative defaults apply when they are omitted:
The terminal requires a Polymarket slug or market_id for real execution so
Foresea can independently obtain a fresh market quote; a raw CLOB token ID alone
is insufficient for this safety check.
To enable the read-only scheduled reconciler, generate one high-entropy service
token and store the same value as Cloud Run's TRADING_RECONCILIATION_TOKEN and
the GitHub Actions secret of that name. This is an operator token, not a user
credential and not an encryption key. The Trading reconciliation workflow then
calls the hidden endpoint every 15 minutes, bounded by
TRADING_RECONCILIATION_MAX_ORDERS (default 25, hard maximum 100). The job
only fetches the current state of already-submitted venue order IDs; it cannot
place, amend, or cancel an order.
After deploying the trading revision, use the same narrowly scoped reconciliation token to read its non-sensitive configuration report:
The report confirms the configured KMS resource, durable store client, reconciliation-token presence, valid hard caps, live-execution gates, and whether the retired shared encryption key is still present. It does not expose key names, tokens, credentials, or account data. It also cannot prove Cloud KMS IAM, that the GitHub Actions secret matches, or that an exchange account can trade; verify those separately during the invite-only smoke test.
Deploy the TradingOrder index in index.yaml before enabling the scheduler:
Required:
question: forecasting question, such as "Will X happen by date Y?",
"Who will win X?", "What will X be?", or "When will X happen?".Optional:
question_type: binary, multiple_choice, numeric, or date. If omitted,
the model attempts to infer the type.options: answer choices for multiple_choice questions.description: extra context for the question.resolution_criteria: how the question should resolve or be measured.categories: list of topic labels.news_articles: caller-supplied evidence articles. If provided, automatic
evidence retrieval is skipped.attach_evidence: defaults to true. When true and news_articles is empty,
the API fetches current evidence from GDELT, Google News RSS, and Stooq.evidence_top_k: number of evidence articles to attach, capped by the server.market_platform: prediction market venue such as Polymarket, Kalshi,
Manifold, or Metaculus.market_url: URL for the market being analyzed.market_outcome: outcome whose market price is supplied. Defaults to Yes
for binary markets.market_probability: current market-implied probability for
market_outcome. Use 0.42 or 42; the API normalizes percentages.variant: prompt variant. Defaults to variant0_neutral_baseline.created_time, publish_time, resolve_time, days_open: optional
forecasting metadata.openrouter_api_key + openrouter_model: run the forecast on your own model
instead of the server default (see "Bring your own model" below).provider_base_url: optional OpenAI-compatible /chat/completions endpoint to
use with your key/model instead of OpenRouter. Must be public HTTPS.By default /predict runs on the server's hosted model. To use your own:
openrouter_api_key and openrouter_model (e.g.
openai/gpt-4o, anthropic/claude-sonnet-4-5). The request is proxied through
OpenRouter.provider_base_url (e.g.
https://api.openai.com/v1 or https://api.openai.com/v1/chat/completions)
with the matching openrouter_model (here just the provider's model ID, e.g.
gpt-4o) and your key. Foresea normalizes /v1 base URLs to
/v1/chat/completions internally.For safety, provider_base_url must be public HTTPS; loopback, private,
link-local, and cloud-metadata hosts are rejected. In the web app, the sidebar's
"Use your own model" panel exposes the provider, endpoint, key, and model.
SCADS AI already exposes Foresea's default models through an OpenAI-compatible hosted endpoint. Use vLLM only when you need direct control over checkpoint, quantization, throughput, or serving hardware.
Start a local vLLM OpenAI-compatible server:
Then point Foresea at the configured qwen3-32b-vllm model:
For production, run Foresea and vLLM as separate services. Foresea's public
bring-your-own endpoint still requires public HTTPS for provider_base_url;
private or loopback vLLM URLs are intended for trusted server-side config.
question_type: detected or requested type: binary, multiple_choice,
numeric, or date.predicted_answer: "Yes", "No", the top multiple-choice option, or the
median numeric/date estimate.confidence: model confidence as a number from 0 to 1 for binary and
multiple-choice forecasts; null for numeric/date forecasts.options: per-option probabilities for multiple-choice forecasts.range_forecast: p10, p50, p90, and optional unit for numeric/date
forecasts.rationale: model-generated explanation.model_rationale: alias for the model-generated explanation, intended for API
clients.evidence_sources: compact source list with article title, URL, publication
date, and relevance score.evidence_articles: full evidence records attached to the prompt.evidence_error: retrieval error message, or null when evidence retrieval
succeeds.market_analysis: optional comparison against a supplied market price:
market_probability, model_probability, edge, stance, and a short
summary. edge is model_probability - market_probability.src/analyzing_llm_rationale/: packaged inference, provider, validation, and CLI logic.configs/: model and rationale-variant definitions.prompts/: system prompt plus the configured rationale, control, ablation,
and no-evidence prompt variants.scripts/: evaluation, recovery, SHAP, perturbation, plotting, market-data,
and utility scripts.slurm/: HPC launchers for the variant/temperature sweeps.results/: model outputs and run metadata.analysis/: aggregate metric tables and rationale-analysis outputs.paper/: paper figures, Draw.io sources, PDFs, and qualitative case studies.tests/: unit tests for the package and metric parsing.See ARTIFACT_MANIFEST.md for the submission checklist and file-level notes.
Use .[dev] for linting and unit tests. Add .[analysis] when regenerating
plots, metrics tables, or SHAP analyses. Add .[trading] for local exchange
order preview/execution development.
Configured variants live in configs/variants.yaml and map directly to prompt
files under prompts/.
variant0 is the neutral baseline.variant1 through variant8 cover the original rationale attribute prompts.variant9 through variant14 add scratchpad, length-matched, structural, and
combined temporal/credibility controls.variant15_neutral_no_rationale and variant16_no_evidence_neutral support
ablations for rationale and evidence effects.When adding a variant, update configs/variants.yaml, add the prompt file, and
run a bounded smoke test:
PYTHONPATH=src is useful when the repository has not been installed yet or an
older user-local install shadows the working tree.
Run the full suite with Python 3.10+ and the relevant extras installed. The
server, RAG, tracking, and trading tests import optional dependencies from
serve, pipeline, analysis, and trading.
Run the variant 3 pipeline with the packaged CLI:
For a remote OpenAI-compatible provider:
If you do not want to install the package into the environment, invoke it directly:
Useful options:
--variant variant6_step_by_step_reasoning: choose the prompt/output contract.--model qwen2.5-7b-instruct: choose a configured model definition.--temperature 0.7: control generation temperature and output directory.--max-records 10: process only a bounded number of records.--reprocess-nulls: rerun existing rows with predicted_answer = null.--drop-article-text: remove raw article text from prompts before inference.--device auto: select cuda when available, otherwise cpu.verify-results --variant ...: verify completeness, duplicates, malformed rows, and missing IDs.validate-dataset: validate the dataset schema before a run.Foresea has a Karpathy-style autoresearch harness for prompt experiments: edit
one candidate prompt, run a fixed benchmark slice, score one metric, and append
an auditable experiment log. The research surface is
autoresearch/candidate_prompt.txt; agent instructions live in
autoresearch/program.md. The default --model gpt-oss-120b uses the
SCADS-hosted OpenAI-compatible endpoint from configs/models.yaml
(SCADS_AI_API_KEY or SCADS_AI_API_KEY.txt).
Run one candidate experiment:
Compare against a baseline and promote only if the candidate improves:
Each run writes analysis/autoresearch/runs/<run_id>/score.json and appends a
machine-readable row to analysis/autoresearch/experiments.jsonl.
Validate an existing result file:
Regenerate aggregate metrics from results/:
Run the DuckDB SQL analytics suite over the real Metaculus-style dataset and saved model outputs:
This writes a markdown report plus one CSV per query for 10 medium-level SQL problems: model accuracy, best variants, calibration bins, Brier score, consensus/disagreement cases, prompt lift over baseline, temperature sensitivity, overconfident errors, and category difficulty.
Run the LangChain-powered news retrieval wrapper:
The news pipeline uses LangChain for a query-planning step, article
summarization, and embedding-based relevance ranking before inference. Evidence
sources are configurable with --source for the CLI and --evidence-source
when serving the API.
Run or schedule the Prefect DAG for RSS/news fetch, inference, and DuckDB logging:
Regenerate paper figures after metrics are present:
Common runner and verification commands:
python scripts/run_variant.py --variant variant5_key_conditionspython scripts/run_variant.py --variant variant3_reasoning_type --temperature 0.7 --temperature-tag temperature_07python scripts/run_variant.py --variant variant4_credibility --model llama-3.3-70b-instructpython scripts/verify_results.py --variant variant3_reasoning_typepython download_qwen_model.pypython check_local_inference.pyRepo layout:
scripts/: modular runner entrypointslurm/: batch launchersAuditability:
run_metadata_<variant>.json next to the results file.Always run unit tests, linter, and knowledge graph sync before committing:
The included dataset is forecasting_qa_news_metaculus_2025-02-01_to_today.metaculus_frs_format.json.
Model access is configured in configs/models.yaml. Open-weight Qwen models run
locally through Hugging Face; hosted models use OpenAI-compatible endpoints and
require API keys through environment variables or local key files.
Never commit key files or tokens. Large local caches (.cache/, envs/, .venv/)
are intentionally ignored and excluded from source archives.
If this repository supports a publication, cite the artifact with the metadata in
CITATION.cff and cite the upstream datasets/models according to their licenses.