Skip to content

Repository files navigation

ClaudIA — AI Trading Assistant

CI

ClaudIA is a Panel-based trading assistant that gives you a persistent, principle-guided AI for market analysis, strategy work, and human-confirmed order staging. It connects to Interactive Brokers via ibkr_core_mcp and to TradingView Desktop via the tradingview-mcp Node.js sidecar.


Features

  • Conversational IBKR access — positions, P&L, live orders, account summary, market data, backtests, price alerts — all via natural language
  • Execution-triggered P&L — a background listener watches for trade executions (any origin — mobile, TWS, web, API) and refreshes account P&L automatically each time a trade fills; no continuous polling
  • Full trade history — 7-year backfill via IBKR Flex Queries; sync_flex_trades keeps it current; get_trades source='store' queries with no date limit
  • Human-confirmed order staging — ClaudIA proposes trades (equities and futures); you click a button → Touch ID → AppKit colored dialog (green/BUY, red/SELL). The LLM has no order-execution tools. CME Rule 536-B fields auto-added for futures
  • TradingView live integration — reads your active chart, sets symbols/timeframes; every Pine-fenced block ClaudIA emits (pine, pinescript or pine-script, any case) gets a Copy button (real client-side clipboard) and an Inject into TradingView button that sets the Pine Editor source directly
  • Live account dashboard — KPI strip · positions · working orders · realised P&L, polled every 15s, read-only by construction. Note the "Realised today" tile follows IBKR's accounting day, which rolls in the late ET evening, not at midnight — see Live Dashboard
  • External candlestick chart pane — a HoloViews/hvplot chart beside the chat (symbol/period/bar controls, volume subplot), OHLCV from the Drive cache with fetch-on-miss from IBKR; fully independent of the conversation
  • Screenshot analysis — upload any TradingView chart for vision-based analysis (no Desktop required)
  • Principle-guided responses — your personal docs/principles.md is loaded as a system prompt; ClaudIA refuses proposals that violate your rules
  • Persistent memory — all sessions, decisions, and symbol observations stored in SQLite with FTS5 search ("what did I decide about NVDA last month?")
  • GDrive sync — claudia.db and context/principles docs auto-sync to Google Drive; pick up any session from any machine
  • Hot-reload documents — edit context.md or principles.md while a session is open; changes apply from the next message
  • Action bar — IBKR / TradingView / Drive buttons under the chat whose colour is the live state (green up, red down, neutral not configured), plus End Session. A click reconnects: IBKR through a read-only pre-flight that never forces a re-login, TradingView by quitting a portless instance and relaunching it with the debug port, Drive by re-authenticating. Colours re-read every 5s over Panel's websocket; the services are polled every 60s (IBKR's /tickle keepalive interval)
  • System log — a collapsed, terminal-style card under the chat for everything that happens to the session (connectivity, Flex sync, document reloads, gateway/TradingView progress, session end); the chat keeps only the conversation
  • Session reports — auto-generated Markdown report at session end: tools called, decisions, errors, connectivity state

Prerequisites

Dependency Purpose
Python 3.11+ ClaudIA runtime
ibkr_core_mcp IBKR tools, gateway management, SQLite store
Docker Desktop IBKR Client Portal Gateway container
Node.js 18+ tradingview-mcp sidecar
TradingView Desktop (macOS) Live chart integration (optional)

Quick Start

# 1. Clone
git clone <this-repo> && cd claudia_ui

# 2. Python env
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"   # pulls ibkr-core-mcp from PyPI as a declared dependency
# Working on ibkr_core_mcp at the same time? Override the PyPI copy with the local
# checkout — same distribution name, so pip replaces it:
#   pip install -e "../ibkr_core_mcp[scraper]" --config-settings editable_mode=strict
# strict mode is required for mypy to resolve an editable core — see CLAUDE.md Dev Setup

# 3. Environment
cp .env.example .env
# Edit .env — minimum: ANTHROPIC_API_KEY

# 4. Personal documents (git-ignored persona + trading rules)
# With GOOGLE_DRIVE_FOLDER_ID set they download from Drive automatically — skip this.
# Otherwise create them by hand:
touch docs/context.md docs/principles.md   # then write persona / trading rules
chmod 600 docs/context.md docs/principles.md

# 5. TradingView sidecar (optional — skip if using screenshot mode only)
git clone https://github.com/tradesdontlie/tradingview-mcp ~/.tradingview-mcp
cd ~/.tradingview-mcp && npm install && cd -   # pure JS — no build step
./scripts/archive-tv-mcp.sh   # snapshot the working version

# 6. Launch
./start-claudia.sh             # recommended: starts IBKR gateway + ClaudIA
# or:
python -m claudia.panel_app    # ClaudIA only — the IBKR button under the chat starts the gateway

Open http://localhost:8001


TradingView Desktop

For live chart integration, TradingView Desktop must be open with remote debugging enabled. ClaudIA launches it for you when you click the TradingView button under the chat (quitting and relaunching an instance that is running without the port), or manually:

open -a "Trading View" --args --remote-debugging-port=9222

If TradingView is already running without the debug port, use the one-command fix — it quits TV gracefully, relaunches with the debug flag, and waits for CDP to come up:

./scripts/launch-tradingview-debug.sh

If the sidecar breaks after a TradingView or npm update, see docs/tradingview-mcp-recovery.md for the break pattern catalog and recovery steps, including a direct CDP from Python fallback.


Architecture

Panel UI (localhost:8001 — native pn.serve Tornado)
    ↓
claudia/panel_app.py        — pn.serve entry: session lifecycle, the reconnect coroutines, layout root
claudia/panel_system_log.py — System log: collapsed terminal-style card for session-level events
claudia/panel_action_bar.py — action bar: IBKR/TradingView/Drive reconnect buttons lit by ConnectivityChecker
claudia/panel_sink.py       — PanelMessageSink: agent output → pn.chat.ChatInterface
claudia/panel_order_flow.py — order/cancel/modify proposal buttons → order_flow cores
claudia/panel_pinescript.py — Pine-fence copy (real clipboard) / inject buttons
claudia/panel_chart.py      — external HoloViews candlestick chart pane (STK, cache-backed)
claudia/agent.py            — Anthropic SDK streaming loop, tool routing, prompt caching
claudia/proposal_tools.py   — strict-schema propose_order/cancel/modify declarations (no execution)
claudia/message_sink.py     — MessageSink protocol (the UI-decoupling seam)
claudia/order_flow.py       — framework-agnostic order-execution cores → biometric gates
claudia/opening_status.py   — UI-free opening-status builders
claudia/context_loader.py   — docs/context.md + docs/principles.md → system prompt
claudia/conversation_store.py — SQLite: sessions, messages, decisions, doc_versions
claudia/status.py           — ConnectivityChecker: polls IBKR/GDrive/TV every 60s
claudia/execution_listener.py — ExecutionListener: WS trade-execution listener; reports every fill to each session (IBKR-authored chat message + toast), then triggers the P&L check
claudia/tradingview.py      — tradingview-mcp sidecar, CDP health, TradingViewBridge
claudia/gdrive_sync.py      — claudia.db + context/principles sync to Google Drive
claudia/session_reporter.py — auto-generate session report at session end
    ↓                               ↓
ibkr_core_mcp               tradingview-mcp (Node.js, localhost stdio)
(local editable install)            ↓
    ↓                       TradingView Desktop (CDP, localhost:9222)
IBKR Client Portal Gateway
(Docker, localhost:5055)

Order Staging

ClaudIA proposes trades; you approve them through two physical gates. Proposing is a tool call that records a proposal and returns a result — it reaches no IBKR API. The LLM has no order-execution tools.

Three properties hold that line, and since 2026-09-14 each fails a test when it stops being true rather than resting on nobody having written the change that breaks it: the model's own layer names no order-write method and imports neither execution module; an execution core is reachable only from a function on_click registers, each of which claims a one-shot before doing anything else; and nothing in the package writes the parameter Panel watches, which is how a button gets pressed without a person. See docs/security-architecture.md § 6.1.

ClaudIA calls propose_order (strict-schema tool — records, executes nothing)
    ↓ agent.py hands the validated input → MessageSink.send_order_proposal()
    ↓ panel_order_flow.render_order_proposal() → button: "STAGE ORDER"
    ↓ User clicks → order_flow._execute_staged_order_core()
    ↓ Gate 1 — Touch ID (macOS LocalAuthentication)
    ↓ Gate 2 — AppKit dialog: green=BUY / red=SELL, 60s auto-cancel, Enter disabled
    ↓ IBKRClient.place_order() → IBKR gateway

Supported instruments:

sec_type Conid resolution Extra fields
STK (default) none — a conid must already be in the proposal (from get_market_snapshot / preview_order, which route through the authoritative resolver) —
FUT /trsrv/futures → front month manualIndicator: True (CME Rule 536-B). extOperator is deliberately not sent — IBKR rejects it with error 8089 despite the docs marking it "Required*"; live-proven 2026-07-24
FOP pre-resolved conid required in the proposal (via get_option_chain) same 536-B fields as FUT

Placement resolves no symbols. The order path used to fall back to /iserver/secdef/search → contracts[0], which carries no country and no currency and has an undocumented result order — so contracts[0] for IGV was the Mexican listing, priced in MXN, in a USD account. FUT is the one exception, resolved by front month and unambiguous by construction. Do not restore symbol resolution here: it would put a second, drifting definition beside the authoritative one.

What the dialog shows is what the body carries. The futures notional is price × qty × multiplier with the multiplier read from /trsrv/futures. Prices render to their own precision rather than to two decimals — most of what this account trades does not tick in cents, and rounding made two prices a full tick apart read identically until 2026-09-14. The card and the dialog share one formatter so they cannot disagree, the card is a deep copy taken before anything is drawn from it, and an order type that needs a price is refused without one instead of being labelled MARKET.

Full field spec and immutability rule: docs/order-api-reference.md, summarized in CLAUDE.md § Order Staging.


Live Dashboard — and the two realised figures

A polling dashboard sits beside the chat: a KPI strip over tabs for Chart, Positions, working Orders, Fills and P&L, refreshed every 15s. Every table is read-only with no click handler bound — cancelling an order stays behind propose_cancel and both gates.

Every dated window on the P&L tab — week, month, YTD — is shown to date: the part IBKR has settled on a Flex statement plus the part not yet on one, both named, and the curve ends at that figure (operator rule 2026-09-29: one rule for all three, "flex realised + realised new"). The tab still shows two realised figures that will not reconcile, and are not meant to — the Realised today tile and the windows are different quantities; never add them or "fix" one to match the other. The dashboard states the reasons on the surface itself:

Realised today (tile) Week / month / YTD (to date)
Source IBKR ledger realizedpnl — live Flex statements (T+1) for the settled part, plus the executions not yet on a statement, reconstructed FIFO from the account's own fills, for the pending part — each part named on the tab
Cost basis IBKR's real-time avgCost the statement basis for the settled part; the traded prices for the pending part
Day boundary IBKR's accounting roll — see below IBKR's session date for the settled part (18:00 ET futures, 20:00 ET stock, 17:00 ET FX); the pending part is "since the last statement" and is never bucketed by a date

The "today" on that tile ends in the late evening, not at midnight

Measured 2026-08-05: the ledger accumulator rolls late in the ET evening, at an hour that varies. Not midnight ET, and not midnight UTC. So for the last hours of a calendar day the tile labelled "Realised today" can already be showing tomorrow — typically 0.00, right after a day that realised something. That reset is correct behaviour, not a data fault, and it is worth knowing before you read the number at 22:30 and conclude the feed broke. This project documented it as a calendar day until that measurement.

Three readings killed three candidates in turn — midnight ET, midnight UTC, then a fixed clock hour (a 37-read watch ending exactly on the surviving bracket's upper bound found the field unmoved). What actually triggers it is a broker-side accounting run, and this project does not know its schedule. No specific time is quoted here on purpose: a disclosure you can check against the clock is worth nothing if it is wrong.

Two further conventions of the same field, both measured to the cent the same evening:

  • It is futures-realised at the traded prices — entry → exit, not settlement-relative. A lot opened at 80.84 and closed the next trade date across a 75.77 settlement still reports against 80.84. Daily variation margin moves cash; it does not re-base this number.
  • Realisation is booked at the closing fill, carrying both legs' commissions.

Mechanics, evidence and the negative controls: read the module docstring of claudia/dashboard_data.py first — it carries the source table for every figure on the screen — then REALISED_LEDGER_WINDOW in the same file.


Agent Behavior — how ClaudIA is stopped from asserting what it did not do

A trading assistant that invents a price, a position or an action it never took is worse than one that says "I don't know". Constraints against that sit in three layers, and which layer a rule lives in decides whether it is a control or a wish:

Layer Where Model can ignore it?
Prompt — user documents context.md / principles.md (Drive) Yes — persona and trading judgment only
Prompt — safety block _SAFETY_BLOCK, agent.py — appended last, non-overridable Yes
Code — evidence four claim detectors + a role:"system" operator channel No
Code — transcript (2026-09-11) text blocks joined with a paragraph break; a contradicted turn withdrawn from the replay; a zero-tool narration retried once before display No

The instruction layer is provably not enough. An audit of the entire conversation store (2026-08-12, all 225 assistant messages) found 23 verified instances across nine sessions since 2026-06-24 where ClaudIA asserted an action or result nothing had produced — and the worst was a fabricated "raw tool result" produced on demand when the user asked for one as an audit, on the same day the prompt rule forbidding exactly that was added. So the rules are backed in code.

The architecture, in one line: the trigger is textual, the verdict is evidence. A detector keys on what the model wrote; the ruling comes from persisted rows and the turn's real tool set — never from asking the model whether it was telling the truth.

Detector Asserts
_claims_completed_proposal claimed a staged order ⇒ a proposal was really recorded
_claims_fresh_book_check claimed a live-book check ⇒ a book tool really ran this turn
_claims_completed_action reported any completed action ⇒ some tool really ran this turn
_claims_verbatim_tool_result showed a "raw tool result" ⇒ some tool really ran this turn

A fire on a turn that ran no tool is withdrawn before display and the turn retried once (2026-09-11: the attempt stays in the store and the decision log, a System-log line is the visible trace, and the retry's first request forces a tool call where the model supports it); a second fire, or a fire on a turn that ran a tool, produces an unhedged correction to the user, a persisted row, its own decision type, and a non-forgeable role:"system" note so the false claim cannot become precedent — and the contradicted row is never replayed as the model's words again. Precision is a measurement, not an argument: frozen at 21 fires (all individually verified fabrications), 22 near-identical texts cleared by their real tool calls, 0 false positives on that corpus — tests/test_corpus_precision.py.

Precision is not recall, and the 2026-09-13 audit measured the other side. The detectors key on four textual shapes; a claim phrased outside them is not detected, and a turn that ran any tool clears two of the four. So this is detection over a named set, not prevention — the record of what really happened is always complete, and the blocking is narrow. The limits are listed in docs/security-architecture.md § 9 rather than left implied.

Of Anthropic's seven documented hallucination techniques (three basic, four advanced) ClaudIA implements four, treats one as analogous, declines one with a reason, and finds one not applicable — the technique-by-technique map, the per-layer rationale and the known limits are in docs/agent-behavior-reference.md.

Agent — Prompt & Context Handling

claudia/agent.py assembles four kinds of information into every API call: the system prompt (context.md + principles.md + market calendar + hardcoded safety block), tool schemas (ClaudeToolkit + TradingView + local tools), conversation history (ConversationStore), and tool results returned mid-loop.

The model is a knob, not a constant. CLAUDIA_MODEL selects the model (claude-opus-4-8 as of 2026-09-11 — a good starting product at a reasonable token cost, and not set in stone: models evolve and other ones will be explored). So nothing in the agent may hardcode an assumption about one model. Every model-dependent behaviour goes through a per-model capability table with a live probe behind each entry, in both directions — the pattern of _OPERATOR_CHANNEL_MODELS / warn_if_model_lacks_operator_channel (the mid-conversation role:"system" channel: accepted on Opus 4.8, 400 on Sonnet 4.6, both probed), and of the forced-tool_choice probe (any under adaptive thinking: accepted on Opus 4.8 with a real tool_use returned, 400 on Fable 5.1 as documented — both probed 2026-09-11, test_live_api_forced_tool_choice_under_adaptive_thinking). Where a model lacks a channel the agent degrades explicitly and loudly — never a silent fallback to a weaker mechanism — and a model that stops rejecting something is a signal to widen the table, not a defect. The constraints a replacement model must satisfy are listed in docs/env-vars-reference.md § CLAUDIA_MODEL.

System prompt — built once per session, not per message. Doc-version and document checks run when ClaudIA loads; a watchdog-driven reload counter (ContextLoader.reload_count) triggers a rebuild only when context.md or principles.md actually changes. Steady-state per-message cost is one integer comparison — no file reads, no DB query.

Prompt caching — 3 breakpoints (cache_control: ephemeral, prefix hierarchy tools → system → messages):

Breakpoint Caches
Last tool definition All tool schemas (42+ IBKR/TV/local tools)
System prompt (block form) Context, principles, calendar, safety block
Last message content block Conversation history, refreshed per API call

Live-verified: a ~22K-token static prefix drops to 0.1× cost on every warm call (vs. full price uncached) — ~90% input-token cost reduction on cached calls. Cache health is logged on every call (prompt cache: created=… read=… uncached=…).

No dead memory tables. sessions, messages, decisions, doc_versions are the only tables — a relationships table and a decisions FTS index were removed 2026-07-03 (never wired to any tool or caller).

Full information-flow map (prompts, session archive, scrape access, and the design constraints a future RAG layer must respect) — docs/audits/2026-07-03-agent-info-architecture-review.md. Implementation plan and live-verified numbers — docs/plans/2026-07-03-prompt-caching-upgrade.md.


Security

ClaudIA can reach a live brokerage account, so the interesting question is not "is it secure" but "which properties are enforced, and which are merely true today". Both are written down.

  • SECURITY.md — the control inventory and how to report a vulnerability
  • docs/security-architecture.md — the living design: principals and what each is not trusted for, the twelve invariants with the test that fails when each stops being true, a dated decision log, and the known limits stated plainly
  • docs/audits/ — point-in-time evidence, never edited after publication

The gates themselves, the order endpoints and the tool capability registry belong to ibkr_core_mcp and are documented there. This project points at them rather than restating them: a copied gate policy here was wrong for three days after the core changed, which is exactly the failure mode the 2026-09-13 audit was looking for.


Documentation

File Contents
docs/README.md Full documentation catalog — every doc in docs/, categorized
CLAUDE.md Developer guide: setup, env vars, architecture, hard rules
SECURITY.md Security model: order barriers, threat model, audit checklist
docs/agent-behavior-reference.md Agent behavior: the safety block, the three enforcement layers, the four claim detectors, and the Anthropic-technique map
docs/flex-query-setup.md IBKR Flex Query setup: token, query config, backfill, ongoing sync
docs/tradingview-mcp-recovery.md TradingView break patterns, recovery steps, CDP fallback
docs/connectivity.md IBKR / GDrive / TradingView check logic, reconnection flows, live test results
docs/project-status.md Milestone history, test coverage, live testing index and log, known gaps

External API Reference

Any contribution touching API behavior, error codes, endpoint paths, or field names must reference the official documentation first — never assume from memory.

API Used in Official reference
IBKR Client Portal API ibkr_core_mcp https://ibkrcampus.com/docs/web-api/
IBKR Flex Web Service ibkr_core_mcp/flex_query.py https://www.ibkrguides.com/clientportal/performanceandstatements/flex3.htm
IBKR Flex error codes ibkr_core_mcp/flex_query.py https://www.ibkrguides.com/clientportal/performanceandstatements/flex3error.htm
Anthropic Messages API claudia/agent.py https://docs.anthropic.com/en/api/messages
Anthropic tool use claudia/agent.py https://docs.anthropic.com/en/docs/build-with-claude/tool-use
Google Drive API v3 claudia/gdrive_sync.py https://developers.google.com/drive/api/reference/rest/v3
TradingView MCP claudia/tradingview.py https://github.com/tradesdontlie/tradingview-mcp
Chrome DevTools Protocol claudia/tradingview.py https://chromedevtools.github.io/devtools-protocol/
Panel claudia/panel_*.py https://panel.holoviz.org
hvPlot / HoloViews claudia/panel_chart.py https://hvplot.holoviz.org / https://holoviews.org
Bokeh (HoloViews' rendering backend) claudia/panel_chart.py https://docs.bokeh.org
requests (web fetch) claudia/agent.py https://docs.python-requests.org/
html2text (HTML → Markdown) claudia/agent.py https://github.com/Alir3z4/html2text
watchdog (file monitoring) claudia/context_loader.py https://watchdog.readthedocs.io/
mcp Python client (stdio) claudia/tradingview.py https://github.com/modelcontextprotocol/python-sdk

Full protocol and per-file ownership: CLAUDE.md → API Reference.


Data Stores

Store Path Contents
claudia.db data/claudia.db Sessions, messages, decisions, doc versions
store.db ~/.ibkr_core/store.db Trade history (Flex), position snapshots, backtests, alerts

Both databases are excluded from git. Run PRAGMA integrity_check to audit health.


Google Drive Architecture (multi-machine portability)

ClaudIA is designed to run on any machine — all persistent state lives in a single Google Drive root folder. Set GOOGLE_DRIVE_FOLDER_ID and ClaudIA restores itself automatically.

<GOOGLE_DRIVE_FOLDER_ID>/          ← one root folder, one env var
  context.md                       ← ClaudIA persona (cloud-authoritative)
  principles.md                    ← trading rules (cloud-authoritative)
  db/
    claudia.db                     ← conversation history (download at start, upload at end)
  market_data/
    manifest.json
    QQQ_1D_6M_2026-06-26.parquet   ← OHLCV cache (shared across machines)
  account_data/
    flex_U123_2026-06-26_REF.xml   ← Flex XML archives (re-importable to SQLite)
    store.db                       ← ibkr_core_mcp trade store backup

What syncs automatically:

Data Direction When
claudia.db Drive → local Session start (first session per process, before DB opens)
claudia.db local → Drive Session end (WAL-consistent backup snapshot — never the live file)
context.md + principles.md Drive → memory Every session start
Flex XML local → account_data/ After every successful Flex sync
OHLCV parquet local → market_data/ After every fetch_market_data call

What a new machine needs (nothing else):

  • GOOGLE_DRIVE_FOLDER_ID — root folder ID
  • GDRIVE_TOKEN_FILE + GDRIVE_CREDENTIALS_FILE — OAuth2 credentials
  • ANTHROPIC_API_KEY — Claude API key
  • IBKR_FLEX_TOKEN + IBKR_FLEX_QUERY_ID — to re-sync trade history from IBKR

store.db is rebuilt from Flex XML archives in account_data/ via sync_flex_archive — no manual export needed.


Testing

pytest                                        # full suite — all unit, no IBKR gateway needed (1,906 collected 2026-09-14)
pytest tests/security                         # the structural invariants, ~3 seconds
ruff check . && ruff format --check . && mypy   # lint, format, type gates

CLAUDIA_LIVE_SCHEMA_CHECK=1 pytest -m live_api   # opt-in; bills real Anthropic API calls

No unit test opens a socket, resolves a name, or sees a real secret. pytest-socket is armed before collection and load_dotenv is neutralised there too, because importing the app loads .env at module scope. Four tests are exempted by name for real DNS, with the reason recorded; one binds a loopback socket because socket behaviour is what it tests. An opted-in live_api run keeps its key and its network, and everything else in that run stays blocked — a block that silently retired those four tests would be worse than no block.

CI adds two gates beyond the four above: a dependency audit of the resolved tree, and a secret scan. If the secret scan is red, read the log before assuming a leak — a failure to run the scanner and a finding look identical in the summary.

The live_api tests are skipped by default. They exist because a local schema validator cannot prove the API accepts a request — three defects that would have returned a 400 on every call once passed both a documentation review and a green suite. Probe the live API before adding a JSON Schema keyword or changing a message-role placement.

Live IBKR verification (order staging, gateway flows) is done manually and recorded in docs/project-status.md § Live Test Log.

About

ClaudIA Trading Assistant is a chat bot UI client using ibkr_core_mcp services

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages