ClaudIA is a Panel-based trading assistant that gives you a persistent, principle-guided AI for market analysis, strategy work, and human-confirmed order staging. It connects to Interactive Brokers via ibkr_core_mcp and to TradingView Desktop via the tradingview-mcp Node.js sidecar.
- Conversational IBKR access — positions, P&L, live orders, account summary, market data, backtests, price alerts — all via natural language
- Execution-triggered P&L — a background listener watches for trade executions (any origin — mobile, TWS, web, API) and refreshes account P&L automatically each time a trade fills; no continuous polling
- Full trade history — 7-year backfill via IBKR Flex Queries;
sync_flex_tradeskeeps it current;get_trades source='store'queries with no date limit - Human-confirmed order staging — ClaudIA proposes trades (equities and futures); you click a button → Touch ID → AppKit colored dialog (green/BUY, red/SELL). The LLM has no order-execution tools. CME Rule 536-B fields auto-added for futures
- TradingView live integration — reads your active chart, sets symbols/timeframes; every Pine-fenced block ClaudIA emits (
pine,pinescriptorpine-script, any case) gets a Copy button (real client-side clipboard) and an Inject into TradingView button that sets the Pine Editor source directly - Live account dashboard — KPI strip · positions · working orders · realised P&L, polled every 15s, read-only by construction. Note the "Realised today" tile follows IBKR's accounting day, which rolls in the late ET evening, not at midnight — see Live Dashboard
- External candlestick chart pane — a HoloViews/hvplot chart beside the chat (symbol/period/bar controls, volume subplot), OHLCV from the Drive cache with fetch-on-miss from IBKR; fully independent of the conversation
- Screenshot analysis — upload any TradingView chart for vision-based analysis (no Desktop required)
- Principle-guided responses — your personal
docs/principles.mdis loaded as a system prompt; ClaudIA refuses proposals that violate your rules - Persistent memory — all sessions, decisions, and symbol observations stored in SQLite with FTS5 search ("what did I decide about NVDA last month?")
- GDrive sync —
claudia.dband context/principles docs auto-sync to Google Drive; pick up any session from any machine - Hot-reload documents — edit
context.mdorprinciples.mdwhile a session is open; changes apply from the next message - Action bar — IBKR / TradingView / Drive buttons under the chat whose colour is the live state (green up, red down, neutral not configured), plus End Session. A click reconnects: IBKR through a read-only pre-flight that never forces a re-login, TradingView by quitting a portless instance and relaunching it with the debug port, Drive by re-authenticating. Colours re-read every 5s over Panel's websocket; the services are polled every 60s (IBKR's
/ticklekeepalive interval) - System log — a collapsed, terminal-style card under the chat for everything that happens to the session (connectivity, Flex sync, document reloads, gateway/TradingView progress, session end); the chat keeps only the conversation
- Session reports — auto-generated Markdown report at session end: tools called, decisions, errors, connectivity state
| Dependency | Purpose |
|---|---|
| Python 3.11+ | ClaudIA runtime |
ibkr_core_mcp |
IBKR tools, gateway management, SQLite store |
| Docker Desktop | IBKR Client Portal Gateway container |
| Node.js 18+ | tradingview-mcp sidecar |
| TradingView Desktop (macOS) | Live chart integration (optional) |
# 1. Clone
git clone <this-repo> && cd claudia_ui
# 2. Python env
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]" # pulls ibkr-core-mcp from PyPI as a declared dependency
# Working on ibkr_core_mcp at the same time? Override the PyPI copy with the local
# checkout — same distribution name, so pip replaces it:
# pip install -e "../ibkr_core_mcp[scraper]" --config-settings editable_mode=strict
# strict mode is required for mypy to resolve an editable core — see CLAUDE.md Dev Setup
# 3. Environment
cp .env.example .env
# Edit .env — minimum: ANTHROPIC_API_KEY
# 4. Personal documents (git-ignored persona + trading rules)
# With GOOGLE_DRIVE_FOLDER_ID set they download from Drive automatically — skip this.
# Otherwise create them by hand:
touch docs/context.md docs/principles.md # then write persona / trading rules
chmod 600 docs/context.md docs/principles.md
# 5. TradingView sidecar (optional — skip if using screenshot mode only)
git clone https://github.com/tradesdontlie/tradingview-mcp ~/.tradingview-mcp
cd ~/.tradingview-mcp && npm install && cd - # pure JS — no build step
./scripts/archive-tv-mcp.sh # snapshot the working version
# 6. Launch
./start-claudia.sh # recommended: starts IBKR gateway + ClaudIA
# or:
python -m claudia.panel_app # ClaudIA only — the IBKR button under the chat starts the gatewayFor live chart integration, TradingView Desktop must be open with remote debugging enabled. ClaudIA launches it for you when you click the TradingView button under the chat (quitting and relaunching an instance that is running without the port), or manually:
open -a "Trading View" --args --remote-debugging-port=9222If TradingView is already running without the debug port, use the one-command fix — it quits TV gracefully, relaunches with the debug flag, and waits for CDP to come up:
./scripts/launch-tradingview-debug.shIf the sidecar breaks after a TradingView or npm update, see
docs/tradingview-mcp-recovery.md for the break pattern catalog
and recovery steps, including a direct CDP from Python fallback.
Panel UI (localhost:8001 — native pn.serve Tornado)
↓
claudia/panel_app.py — pn.serve entry: session lifecycle, the reconnect coroutines, layout root
claudia/panel_system_log.py — System log: collapsed terminal-style card for session-level events
claudia/panel_action_bar.py — action bar: IBKR/TradingView/Drive reconnect buttons lit by ConnectivityChecker
claudia/panel_sink.py — PanelMessageSink: agent output → pn.chat.ChatInterface
claudia/panel_order_flow.py — order/cancel/modify proposal buttons → order_flow cores
claudia/panel_pinescript.py — Pine-fence copy (real clipboard) / inject buttons
claudia/panel_chart.py — external HoloViews candlestick chart pane (STK, cache-backed)
claudia/agent.py — Anthropic SDK streaming loop, tool routing, prompt caching
claudia/proposal_tools.py — strict-schema propose_order/cancel/modify declarations (no execution)
claudia/message_sink.py — MessageSink protocol (the UI-decoupling seam)
claudia/order_flow.py — framework-agnostic order-execution cores → biometric gates
claudia/opening_status.py — UI-free opening-status builders
claudia/context_loader.py — docs/context.md + docs/principles.md → system prompt
claudia/conversation_store.py — SQLite: sessions, messages, decisions, doc_versions
claudia/status.py — ConnectivityChecker: polls IBKR/GDrive/TV every 60s
claudia/execution_listener.py — ExecutionListener: WS trade-execution listener; reports every fill to each session (IBKR-authored chat message + toast), then triggers the P&L check
claudia/tradingview.py — tradingview-mcp sidecar, CDP health, TradingViewBridge
claudia/gdrive_sync.py — claudia.db + context/principles sync to Google Drive
claudia/session_reporter.py — auto-generate session report at session end
↓ ↓
ibkr_core_mcp tradingview-mcp (Node.js, localhost stdio)
(local editable install) ↓
↓ TradingView Desktop (CDP, localhost:9222)
IBKR Client Portal Gateway
(Docker, localhost:5055)
ClaudIA proposes trades; you approve them through two physical gates. Proposing is a tool call that records a proposal and returns a result — it reaches no IBKR API. The LLM has no order-execution tools.
Three properties hold that line, and since 2026-09-14 each fails a test when it stops being
true rather than resting on nobody having written the change that breaks it: the model's own
layer names no order-write method and imports neither execution module; an execution core is
reachable only from a function on_click registers, each of which claims a one-shot before
doing anything else; and nothing in the package writes the parameter Panel watches, which is
how a button gets pressed without a person. See
docs/security-architecture.md § 6.1.
ClaudIA calls propose_order (strict-schema tool — records, executes nothing)
↓ agent.py hands the validated input → MessageSink.send_order_proposal()
↓ panel_order_flow.render_order_proposal() → button: "STAGE ORDER"
↓ User clicks → order_flow._execute_staged_order_core()
↓ Gate 1 — Touch ID (macOS LocalAuthentication)
↓ Gate 2 — AppKit dialog: green=BUY / red=SELL, 60s auto-cancel, Enter disabled
↓ IBKRClient.place_order() → IBKR gateway
Supported instruments:
sec_type |
Conid resolution | Extra fields |
|---|---|---|
STK (default) |
none — a conid must already be in the proposal (from get_market_snapshot / preview_order, which route through the authoritative resolver) |
— |
FUT |
/trsrv/futures → front month |
manualIndicator: True (CME Rule 536-B). extOperator is deliberately not sent — IBKR rejects it with error 8089 despite the docs marking it "Required*"; live-proven 2026-07-24 |
FOP |
pre-resolved conid required in the proposal (via get_option_chain) |
same 536-B fields as FUT |
Placement resolves no symbols. The order path used to fall back to
/iserver/secdef/search → contracts[0], which carries no country and no currency and
has an undocumented result order — so contracts[0] for IGV was the Mexican
listing, priced in MXN, in a USD account. FUT is the one exception, resolved by front
month and unambiguous by construction. Do not restore symbol resolution here: it would
put a second, drifting definition beside the authoritative one.
What the dialog shows is what the body carries. The futures notional is
price × qty × multiplier with the multiplier read from /trsrv/futures. Prices render to
their own precision rather than to two decimals — most of what this account trades does not
tick in cents, and rounding made two prices a full tick apart read identically until
2026-09-14. The card and the dialog share one formatter so they cannot disagree, the card is
a deep copy taken before anything is drawn from it, and an order type that needs a price is
refused without one instead of being labelled MARKET.
Full field spec and immutability rule: docs/order-api-reference.md,
summarized in CLAUDE.md § Order Staging.
A polling dashboard sits beside the chat: a KPI strip over tabs for Chart, Positions,
working Orders, Fills and P&L, refreshed every 15s. Every table is read-only with no click
handler bound — cancelling an order stays behind propose_cancel and both gates.
Every dated window on the P&L tab — week, month, YTD — is shown to date: the part IBKR has settled on a Flex statement plus the part not yet on one, both named, and the curve ends at that figure (operator rule 2026-09-29: one rule for all three, "flex realised + realised new"). The tab still shows two realised figures that will not reconcile, and are not meant to — the Realised today tile and the windows are different quantities; never add them or "fix" one to match the other. The dashboard states the reasons on the surface itself:
| Realised today (tile) | Week / month / YTD (to date) | |
|---|---|---|
| Source | IBKR ledger realizedpnl — live |
Flex statements (T+1) for the settled part, plus the executions not yet on a statement, reconstructed FIFO from the account's own fills, for the pending part — each part named on the tab |
| Cost basis | IBKR's real-time avgCost |
the statement basis for the settled part; the traded prices for the pending part |
| Day boundary | IBKR's accounting roll — see below | IBKR's session date for the settled part (18:00 ET futures, 20:00 ET stock, 17:00 ET FX); the pending part is "since the last statement" and is never bucketed by a date |
Measured 2026-08-05: the ledger accumulator rolls late in the ET evening, at an hour that varies. Not midnight ET, and not midnight UTC. So for the last hours of a calendar day the tile labelled "Realised today" can already be showing tomorrow — typically 0.00, right after a day that realised something. That reset is correct behaviour, not a data fault, and it is worth knowing before you read the number at 22:30 and conclude the feed broke. This project documented it as a calendar day until that measurement.
Three readings killed three candidates in turn — midnight ET, midnight UTC, then a fixed clock hour (a 37-read watch ending exactly on the surviving bracket's upper bound found the field unmoved). What actually triggers it is a broker-side accounting run, and this project does not know its schedule. No specific time is quoted here on purpose: a disclosure you can check against the clock is worth nothing if it is wrong.
Two further conventions of the same field, both measured to the cent the same evening:
- It is futures-realised at the traded prices — entry → exit, not settlement-relative. A lot opened at 80.84 and closed the next trade date across a 75.77 settlement still reports against 80.84. Daily variation margin moves cash; it does not re-base this number.
- Realisation is booked at the closing fill, carrying both legs' commissions.
Mechanics, evidence and the negative controls: read the module docstring of
claudia/dashboard_data.py first — it carries the source
table for every figure on the screen — then REALISED_LEDGER_WINDOW in the same file.
A trading assistant that invents a price, a position or an action it never took is worse than one that says "I don't know". Constraints against that sit in three layers, and which layer a rule lives in decides whether it is a control or a wish:
| Layer | Where | Model can ignore it? |
|---|---|---|
| Prompt — user documents | context.md / principles.md (Drive) |
Yes — persona and trading judgment only |
| Prompt — safety block | _SAFETY_BLOCK, agent.py — appended last, non-overridable |
Yes |
| Code — evidence | four claim detectors + a role:"system" operator channel |
No |
| Code — transcript (2026-09-11) | text blocks joined with a paragraph break; a contradicted turn withdrawn from the replay; a zero-tool narration retried once before display | No |
The instruction layer is provably not enough. An audit of the entire conversation store (2026-08-12, all 225 assistant messages) found 23 verified instances across nine sessions since 2026-06-24 where ClaudIA asserted an action or result nothing had produced — and the worst was a fabricated "raw tool result" produced on demand when the user asked for one as an audit, on the same day the prompt rule forbidding exactly that was added. So the rules are backed in code.
The architecture, in one line: the trigger is textual, the verdict is evidence. A detector keys on what the model wrote; the ruling comes from persisted rows and the turn's real tool set — never from asking the model whether it was telling the truth.
| Detector | Asserts |
|---|---|
_claims_completed_proposal |
claimed a staged order ⇒ a proposal was really recorded |
_claims_fresh_book_check |
claimed a live-book check ⇒ a book tool really ran this turn |
_claims_completed_action |
reported any completed action ⇒ some tool really ran this turn |
_claims_verbatim_tool_result |
showed a "raw tool result" ⇒ some tool really ran this turn |
A fire on a turn that ran no tool is withdrawn before display and the turn retried once
(2026-09-11: the attempt stays in the store and the decision log, a System-log line is the
visible trace, and the retry's first request forces a tool call where the model supports it);
a second fire, or a fire on a turn that ran a tool, produces an unhedged correction to the
user, a persisted row, its own decision type, and a non-forgeable role:"system" note so the
false claim cannot become precedent — and the contradicted row is never replayed as the
model's words again. Precision is a measurement, not an argument: frozen at 21 fires (all individually
verified fabrications), 22 near-identical texts cleared by their real tool calls, 0 false
positives on that corpus — tests/test_corpus_precision.py.
Precision is not recall, and the 2026-09-13 audit measured the other side. The detectors
key on four textual shapes; a claim phrased outside them is not detected, and a turn that ran
any tool clears two of the four. So this is detection over a named set, not prevention — the
record of what really happened is always complete, and the blocking is narrow. The limits are
listed in docs/security-architecture.md § 9 rather than
left implied.
Of Anthropic's seven documented hallucination techniques
(three basic, four advanced) ClaudIA implements four, treats one as analogous, declines
one with a reason, and finds one not applicable — the technique-by-technique map, the
per-layer rationale and the known limits are in
docs/agent-behavior-reference.md.
claudia/agent.py assembles four kinds of information into every API call: the
system prompt (context.md + principles.md + market calendar + hardcoded safety
block), tool schemas (ClaudeToolkit + TradingView + local tools), conversation
history (ConversationStore), and tool results returned mid-loop.
The model is a knob, not a constant. CLAUDIA_MODEL selects the model (claude-opus-4-8
as of 2026-09-11 — a good starting product at a reasonable token cost, and not set in stone:
models evolve and other ones will be explored). So nothing in the agent may hardcode an
assumption about one model. Every model-dependent behaviour goes through a per-model
capability table with a live probe behind each entry, in both directions — the pattern of
_OPERATOR_CHANNEL_MODELS / warn_if_model_lacks_operator_channel (the mid-conversation
role:"system" channel: accepted on Opus 4.8, 400 on Sonnet 4.6, both probed), and of the
forced-tool_choice probe (any under adaptive thinking: accepted on Opus 4.8 with a real
tool_use returned, 400 on Fable 5.1 as documented — both probed 2026-09-11,
test_live_api_forced_tool_choice_under_adaptive_thinking). Where a model lacks a channel the
agent degrades explicitly and loudly — never a silent fallback to a weaker mechanism — and
a model that stops rejecting something is a signal to widen the table, not a defect. The
constraints a replacement model must satisfy are listed in
docs/env-vars-reference.md § CLAUDIA_MODEL.
System prompt — built once per session, not per message. Doc-version and
document checks run when ClaudIA loads; a watchdog-driven reload counter
(ContextLoader.reload_count) triggers a rebuild only when context.md or
principles.md actually changes. Steady-state per-message cost is one integer
comparison — no file reads, no DB query.
Prompt caching — 3 breakpoints (cache_control: ephemeral, prefix hierarchy
tools → system → messages):
| Breakpoint | Caches |
|---|---|
| Last tool definition | All tool schemas (42+ IBKR/TV/local tools) |
| System prompt (block form) | Context, principles, calendar, safety block |
| Last message content block | Conversation history, refreshed per API call |
Live-verified: a ~22K-token static prefix drops to 0.1× cost on every warm call
(vs. full price uncached) — ~90% input-token cost reduction on cached calls.
Cache health is logged on every call (prompt cache: created=… read=… uncached=…).
No dead memory tables. sessions, messages, decisions, doc_versions are
the only tables — a relationships table and a decisions FTS index were
removed 2026-07-03 (never wired to any tool or caller).
Full information-flow map (prompts, session archive, scrape access, and the
design constraints a future RAG layer must respect) —
docs/audits/2026-07-03-agent-info-architecture-review.md.
Implementation plan and live-verified numbers —
docs/plans/2026-07-03-prompt-caching-upgrade.md.
ClaudIA can reach a live brokerage account, so the interesting question is not "is it secure" but "which properties are enforced, and which are merely true today". Both are written down.
SECURITY.md— the control inventory and how to report a vulnerabilitydocs/security-architecture.md— the living design: principals and what each is not trusted for, the twelve invariants with the test that fails when each stops being true, a dated decision log, and the known limits stated plainlydocs/audits/— point-in-time evidence, never edited after publication
The gates themselves, the order endpoints and the tool capability registry belong to
ibkr_core_mcp and are documented there. This project points at them rather than restating
them: a copied gate policy here was wrong for three days after the core changed, which is
exactly the failure mode the 2026-09-13 audit was looking for.
| File | Contents |
|---|---|
docs/README.md |
Full documentation catalog — every doc in docs/, categorized |
CLAUDE.md |
Developer guide: setup, env vars, architecture, hard rules |
SECURITY.md |
Security model: order barriers, threat model, audit checklist |
docs/agent-behavior-reference.md |
Agent behavior: the safety block, the three enforcement layers, the four claim detectors, and the Anthropic-technique map |
docs/flex-query-setup.md |
IBKR Flex Query setup: token, query config, backfill, ongoing sync |
docs/tradingview-mcp-recovery.md |
TradingView break patterns, recovery steps, CDP fallback |
docs/connectivity.md |
IBKR / GDrive / TradingView check logic, reconnection flows, live test results |
docs/project-status.md |
Milestone history, test coverage, live testing index and log, known gaps |
Any contribution touching API behavior, error codes, endpoint paths, or field names must reference the official documentation first — never assume from memory.
| API | Used in | Official reference |
|---|---|---|
| IBKR Client Portal API | ibkr_core_mcp |
https://ibkrcampus.com/docs/web-api/ |
| IBKR Flex Web Service | ibkr_core_mcp/flex_query.py |
https://www.ibkrguides.com/clientportal/performanceandstatements/flex3.htm |
| IBKR Flex error codes | ibkr_core_mcp/flex_query.py |
https://www.ibkrguides.com/clientportal/performanceandstatements/flex3error.htm |
| Anthropic Messages API | claudia/agent.py |
https://docs.anthropic.com/en/api/messages |
| Anthropic tool use | claudia/agent.py |
https://docs.anthropic.com/en/docs/build-with-claude/tool-use |
| Google Drive API v3 | claudia/gdrive_sync.py |
https://developers.google.com/drive/api/reference/rest/v3 |
| TradingView MCP | claudia/tradingview.py |
https://github.com/tradesdontlie/tradingview-mcp |
| Chrome DevTools Protocol | claudia/tradingview.py |
https://chromedevtools.github.io/devtools-protocol/ |
| Panel | claudia/panel_*.py |
https://panel.holoviz.org |
| hvPlot / HoloViews | claudia/panel_chart.py |
https://hvplot.holoviz.org / https://holoviews.org |
| Bokeh (HoloViews' rendering backend) | claudia/panel_chart.py |
https://docs.bokeh.org |
requests (web fetch) |
claudia/agent.py |
https://docs.python-requests.org/ |
html2text (HTML → Markdown) |
claudia/agent.py |
https://github.com/Alir3z4/html2text |
watchdog (file monitoring) |
claudia/context_loader.py |
https://watchdog.readthedocs.io/ |
mcp Python client (stdio) |
claudia/tradingview.py |
https://github.com/modelcontextprotocol/python-sdk |
Full protocol and per-file ownership: CLAUDE.md → API Reference.
| Store | Path | Contents |
|---|---|---|
claudia.db |
data/claudia.db |
Sessions, messages, decisions, doc versions |
store.db |
~/.ibkr_core/store.db |
Trade history (Flex), position snapshots, backtests, alerts |
Both databases are excluded from git. Run PRAGMA integrity_check to audit health.
ClaudIA is designed to run on any machine — all persistent state lives in a single Google Drive root folder. Set GOOGLE_DRIVE_FOLDER_ID and ClaudIA restores itself automatically.
<GOOGLE_DRIVE_FOLDER_ID>/ ← one root folder, one env var
context.md ← ClaudIA persona (cloud-authoritative)
principles.md ← trading rules (cloud-authoritative)
db/
claudia.db ← conversation history (download at start, upload at end)
market_data/
manifest.json
QQQ_1D_6M_2026-06-26.parquet ← OHLCV cache (shared across machines)
account_data/
flex_U123_2026-06-26_REF.xml ← Flex XML archives (re-importable to SQLite)
store.db ← ibkr_core_mcp trade store backup
What syncs automatically:
| Data | Direction | When |
|---|---|---|
claudia.db |
Drive → local | Session start (first session per process, before DB opens) |
claudia.db |
local → Drive | Session end (WAL-consistent backup snapshot — never the live file) |
context.md + principles.md |
Drive → memory | Every session start |
| Flex XML | local → account_data/ |
After every successful Flex sync |
| OHLCV parquet | local → market_data/ |
After every fetch_market_data call |
What a new machine needs (nothing else):
GOOGLE_DRIVE_FOLDER_ID— root folder IDGDRIVE_TOKEN_FILE+GDRIVE_CREDENTIALS_FILE— OAuth2 credentialsANTHROPIC_API_KEY— Claude API keyIBKR_FLEX_TOKEN+IBKR_FLEX_QUERY_ID— to re-sync trade history from IBKR
store.db is rebuilt from Flex XML archives in account_data/ via sync_flex_archive — no manual export needed.
pytest # full suite — all unit, no IBKR gateway needed (1,906 collected 2026-09-14)
pytest tests/security # the structural invariants, ~3 seconds
ruff check . && ruff format --check . && mypy # lint, format, type gates
CLAUDIA_LIVE_SCHEMA_CHECK=1 pytest -m live_api # opt-in; bills real Anthropic API callsNo unit test opens a socket, resolves a name, or sees a real secret. pytest-socket is
armed before collection and load_dotenv is neutralised there too, because importing the app
loads .env at module scope. Four tests are exempted by name for real DNS, with the reason
recorded; one binds a loopback socket because socket behaviour is what it tests. An opted-in
live_api run keeps its key and its network, and everything else in that run stays blocked —
a block that silently retired those four tests would be worse than no block.
CI adds two gates beyond the four above: a dependency audit of the resolved tree, and a secret scan. If the secret scan is red, read the log before assuming a leak — a failure to run the scanner and a finding look identical in the summary.
The live_api tests are skipped by default. They exist because a local schema validator
cannot prove the API accepts a request — three defects that would have returned a 400 on
every call once passed both a documentation review and a green suite. Probe the live API
before adding a JSON Schema keyword or changing a message-role placement.
Live IBKR verification (order staging, gateway flows) is done manually and recorded in
docs/project-status.md § Live Test Log.