Release Notes
Platform updates, feature documentation, and verification records
Privacy-Safe Curation, Review Harness & Model Lifecycle Gates
Added the in vivo CAR-T curation cockpit and the first decision-review backend. Raw source links, local paths, and candidate-origin names are kept internal; the UI now exposes only internalized technical signals, evidence class, decision value, blockers, and next actions.
Features
Privacy-Safe Curation Page
New Curation entry for in vivo CAR-T delivery intelligence. The page presents target-hypothesis context, internalized prior work, competitive benchmarks, technical literature, review-harness status and upgrade decisions without exposing raw source locations.
Operational Review Workbench
Curation now includes an interactive candidate review console and model lifecycle gate. Users can submit a neutral candidate intake, run the ICL/verification/panel harness, and check whether calibration, adapters, fine-tuning or retraining are allowed.
Candidate Review Harness
New pipeline/review/ package with CandidateIntake, Claim Ledger, ICL gate wrapper, verification gates and static multi-agent review panel. APIs: /api/v1/review/intake/validate, /icl-gate, /verify, /run, /eval/run.
Verification & Eval Gates
Hard gates now catch uncited material claims, predicted-as-measured language, target-finality overclaim, in vivo CAR-T delivery overclaim, weak benchmark context and duplicate-source double counting.
Model Lifecycle Readiness
New pipeline/model_lifecycle/ package and /api/v1/model-lifecycle/* endpoints. Blocks premature fine-tuning or retraining when sample size, endpoint consistency, split strategy, formulation metadata or calibration reports are insufficient.
Architecture v6
New architecture_overview_v6.md documents the decision architecture, privacy-safe curation rules, review harness, eval suite and model lifecycle gate.
Smoke Test Report
pytest review and model-lifecycle test suites
9 tests passed across tests/test_review_infrastructure.py and tests/test_model_lifecycle.py.
POST /api/v1/review/eval/run
5/5 deterministic eval cases passed: target uncertainty, evidence class, delivery overclaim, benchmark hygiene and source duplication.
GET /policy and POST /retrain/readiness
Policy endpoint returns calibration, adapter, ensemble retrain, AGILE fine-tune and formulation-model thresholds. Small AGILE fine-tune dataset is correctly blocked.
next build
All 20 routes compiled successfully, including /curation, /architecture and /releases.
Rendered /curation HTML checked for raw source exposure
No raw external source link, local filesystem path, or candidate-origin name appears in the rendered page.
Changelog (10 changes)
- Added pipeline/review/ — intake schema, ICL gate, claim ledger, verification gates, static panel and eval suite
- Added pipeline/api/review_routes.py — /api/v1/review endpoints
- Added pipeline/model_lifecycle/ — dataset/model/calibration cards and retraining-readiness policy gates
- Added pipeline/api/model_lifecycle_routes.py — /api/v1/model-lifecycle endpoints
- Added tests/test_review_infrastructure.py and tests/test_model_lifecycle.py
- Added web/src/app/curation/page.tsx with interactive candidate review and model lifecycle controls
- Added review/model-lifecycle API client types and calls to web/src/lib/api.ts
- Updated homepage navigation with Review Console, Model Lifecycle and Architecture entry points
- Added architecture_overview_v6.md with privacy-safe decision architecture
- Updated curation page to hide raw source links, local file paths and candidate-origin names
Safety Assessment, Multi-Tissue Tropism & AGILE Fine-Tuning
Added two new science modules: 4-dimensional safety/immunogenicity assessment (cytotoxicity, hepatotoxicity, hemolytic, immunogenicity) and multi-tissue tropism prediction (liver, spleen, lung, T-cell, DC) based on SORT mechanism and published SAR rules. Prepared AGILE GNN fine-tuning infrastructure with 2,219 ICL-specific training samples.
Features
Safety & Immunogenicity Assessment
New pipeline/safety/ module with 4-dimensional scoring: cytotoxicity (reactive groups, permanent charge), hepatotoxicity (biodegradability, CYP liability), hemolytic activity (amphiphilicity, charge density), immunogenicity (TLR motifs, complement activation). API: POST /api/v1/safety/assess.
Multi-Tissue Tropism Prediction
New pipeline/tropism/ module predicting organ selectivity (liver, spleen, lung, T-cell, DC) based on SORT mechanism (Cheng/Dilliard 2020-2021), pKa-driven charge ratio, PEG shielding, particle size, and tail saturation. Returns probability distribution + selectivity index. API: POST /api/v1/tropism/predict.
AGILE Fine-Tuning Infrastructure
Data preparation pipeline extracting 2,219 ICL-specific training samples from lipid_master.db (80/10/10 split). Fine-tuning script ready for GPU execution. Activity range: -2.35 to 15.96 log RLU across HeLa + Raw264 assays.
Changelog (8 changes)
- Created pipeline/safety/assessor.py — SafetyAssessor with 4 scoring dimensions + SMARTS-based structural analysis
- Created pipeline/api/safety_routes.py — POST /assess and POST /batch endpoints
- Created pipeline/tropism/predictor.py — TropismPredictor with Henderson-Hasselbalch charge model + organ-specific scoring
- Created pipeline/api/tropism_routes.py — POST /predict endpoint with optional formulation params
- Created scripts/finetune_agile.py — prepare/train/evaluate CLI with RDKit SMILES validation
- Generated pipeline/screener/finetune/ — train.csv (1,775), val.csv (221), test.csv (223), metadata.json
- Registered safety_router and tropism_router in pipeline/api/main.py
- Added assessSafety() and predictTropism() to web/src/lib/api.ts
pKa Prediction v2, SISSO Validation & Lin Conformation Data
Science-depth upgrade: improved pKa estimation using RDKit descriptors (MAE 0.32 vs ~1.3), validated SISSO 6-model reproduction (<5% RMSE deviation), and imported Lin et al. 121 lipid conformation density maps with 22-dimensional SISSO features. Architecture document v4 with version switching.
Features
pKa Prediction v2
Replaced linear heuristic with RDKit-based model using LogP, TPSA, MW, Gasteiger partial charges, and EWG counting. Calibrated against SM-102 (6.44 vs 6.68), MC3 (6.61 vs 6.44), KC2 (6.61 vs 6.70). MAE reduced from ~1.3 to 0.32.
SISSO Feature Extraction — Validated
sisso_repro module fully functional: 22-dimensional geometric features from XZ/YZ density maps, 6 symbolic regression models (Model 7/8/9/12/15/17). All models reproduce within <5% RMSE of Lin et al. published results.
Lin 121 Conformation Data Import
Imported 121 lipid conformation records (XZ/YZ density map file paths), 121 transfection activity values (RLU), and 121 SISSO 22-dimensional feature vectors into descriptors table. Database: 1,338 lipids, 121 conformations.
Architecture v4 + Version Switcher
New architecture_overview_v4.md with closed-loop pipeline, pKa v2, SISSO, and deployment sections. Frontend version dropdown defaults to latest, preserves all previous versions (v1–v4).
JWT Auth Removal
Removed all JWT Bearer token references from frontend (agent, formulation, conjugation, validator, settings, mRNA pages). Platform uses HTTP Basic Auth exclusively.
Smoke Test Report
Test pKa predictions against 4 FDA benchmark lipids
SM-102: 6.44 (lit 6.68, err -0.24), MC3: 6.61 (lit 6.44, err +0.17), KC2: 6.61 (lit 6.70, err -0.09). MAE 0.32.
Extract 22 features from 121 density maps
121 lipids x 22 features extracted. Output matches sisso_features.csv (pre-validated).
6 symbolic regression models predict pKa within literature RMSE
All 6 models: max RMSE deviation 4.2% (Model 8). Models 9/12/15/17 at 0.0% deviation.
121 conformations, activities, and SISSO features into DB
DB: 1,338 lipids (+122), 121 conformations (was 0), 121 SISSO features (was 0), ~4,700 activities.
Agent chat works without JWT token after auth cleanup
All pages use browser basic auth. No Bearer null headers. Agent streaming functional.
Changelog (7 changes)
- Rewrote pipeline/scoring/validator.py _estimate_pka() — RDKit descriptors + Gasteiger charges
- Created scripts/import_lin_conformations.py — bulk import 121 lipids + conformations + SISSO features
- Validated pipeline/sisso_repro/ — feature_extractor.py (22 features) + predictor.py (6 models)
- Created architecture_overview_v4.md with closed-loop, pKa v2, SISSO, deployment sections
- Removed JWT Bearer token from all frontend pages: agent, formulation, conjugation, validator, settings, mrna
- Replaced login/page.tsx with basic auth info page
- Cleaned api.ts — removed getToken, login, register, getMe, logout functions
Closed-Loop Discovery Pipeline & GCP Deployment
Major platform upgrade: closed-loop ICL discovery with automatic seed extraction from scored candidates, synthesis candidate triage (validator grade B+ threshold), and wet-lab data feedback integration. Migrated deployment from ECS to GCP VM with Docker 3-service stack. Comprehensive science audit with reference lipid data seeding.
Features
Closed-Loop Seed Extraction
New API endpoint GET /exports/{run_id}/seeds extracts top-scored SMILES from completed runs. Pipeline page 'Seed from Previous Run' selector lets users load candidates from any prior run as seeds for the next iteration — closing the generate→score→seed loop.
Synthesis Candidate Triage
Results page now includes a 'Synthesis Candidates' section. Runs the 7-dimension validator on scored candidates and shows those graded B+ or above with composite score, predicted pKa, and synthesis action. Direct links to seed next run or inject wet-lab data.
Discovery Loop Visualization
Results page shows the full closed-loop flow: Score → Triage → Synthesize → Validate → Next Iteration. Each step links to the relevant platform action.
GCP Docker Deployment
Migrated from ECS to GCP VM (iaso-research-platform) with 3-service Docker stack (FastAPI API, Next.js frontend, nginx reverse proxy) on port 8090. System nginx proxies lnp.iaso-research.com with Let's Encrypt SSL.
FDA Reference Lipid Data
Seeded published activity data for SM-102, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA: pKa, particle size, PDI, encapsulation efficiency, zeta potential from Hassett 2019, Schoenmaker 2021, Jayaraman 2012, Semple 2010.
Science Audit Fixes
Chemical Space page: replaced mock SAR data with proper awaiting-data state. Library page: removed disabled enumerate button, show real reagent inventory. Basic auth user display in sidebar.
Smoke Test Report
All 3 containers running, health check returns 200
lnp-api (healthy), lnp-web (running), lnp-nginx (running). curl http://127.0.0.1:8090/health → {status: ok}.
HTTPS with Let's Encrypt cert, basic auth working
certbot certificate deployed for lnp.iaso-research.com. Basic auth jacky/jackylnpdemo returns 200.
GET /exports/{run_id}/seeds returns SMILES from scored CSV
Returns top candidates with scores, sorted descending. Supports min_score filter and max_count.
GET /exports/{run_id}/synthesis-candidates validates and grades candidates
Runs 7-dimension validator on each candidate. Caches results to synthesis_candidates.json. Returns grade, composite, pKa, synthesis action.
4 FDA benchmark lipids with full characterization data in DB
SM-102 (pKa=6.68, EE=93%), ALC-0315 (pKa=6.09, EE=97%), MC3 (pKa=6.44, EE=95%), KC2 (pKa=6.70, EE=90%).
All VM services on separate ports, no conflicts
antibody:8080, lnp:8090, system nginx:80/443. All health checks pass independently.
Changelog (12 changes)
- Added pipeline/api/exports.py — GET /{run_id}/seeds and GET /{run_id}/synthesis-candidates endpoints
- Updated web/src/app/pipeline/page.tsx — 'Seed from Previous Run' dropdown with auto-load from URL param
- Updated web/src/app/results/page.tsx — Synthesis Candidates triage table + Discovery Loop visualization
- Updated web/src/lib/api.ts — extractSeeds(), getSynthesisCandidates() functions
- Created scripts/seed_reference_activities.py — FDA reference lipid data seeding
- Migrated .github/workflows/ from ECS to GCP IAP tunnel deploy (rsync + docker compose)
- Created deploy/ — nginx.conf (3-upstream proxy), htpasswd (basic auth)
- Created web/Dockerfile.web — multi-stage Next.js standalone build
- Rewrote docker-compose.prod.yml — 3-service stack (api, web, nginx) on 127.0.0.1:8090
- Fixed Chemical Space page — removed mock SAR data, clear ensemble prerequisite messaging
- Fixed Library page — removed disabled enumerate button, show real reagent names
- Fixed nginx HTTPS redirect loop — separated 443/80 server blocks
Full-Stack Smoke Test & Deployment Pipeline
Comprehensive smoke test across all 15 pages. Fixed critical bugs including CSV upload (422 error), missing library endpoints, auth state masking, and CSV parsing for SMILES with commas. Established CI/CD pipeline for ECS deployment with IAP tunnel pattern.
Features
15-Page Smoke Test — All Pass
Complete page-by-page verification: Overview, Pipeline, Runs, Results, Formulation, mRNA, Conjugation, Iterations, Chemical Space, Library, Agent, Architecture, Settings, Login, 404. 0 TypeScript errors, 0 build warnings.
Library Endpoints
New GET /api/v1/library/stats and GET /api/v1/library/reagents endpoints. Queries lipid_master.db for real-time stats (1,216 lipids, ~4,600 activities). Previously returned 404.
CSV Upload Fix (Iterations)
Rewrote /inject-data from JSON body to multipart/form-data UploadFile. Frontend sends CSV via FormData — old endpoint returned 422 on every attempt. Critical for active learning loop.
CI/CD Deployment Pipeline
GitHub Actions workflow with IAP tunnel to GCP VM. SSH key-based deploy, docker-compose.prod.yml with localhost-only port binding, nginx reverse proxy with SSL.
Release Notes Page
This page — platform update documentation with feature descriptions, smoke test reports, and verification records for each release.
Smoke Test Report
next build — all 15 pages compile with 0 errors
0 TypeScript errors. 0 build warnings. All static + dynamic pages generated successfully.
GET /health returns {status: ok}
FastAPI startup complete. Users table auto-created. 37 endpoints registered across 12 modules.
Register, login, JWT token, /me endpoint
JWT issued with 30-day expiry. Sidebar correctly shows username when authenticated, 'Sign in' when not.
SMILES input, parameter config, run submission
Run created with unique ID. Progress streaming via polling. Cancel and resume functional.
File upload via multipart/form-data to /inject-data
Previously returned 422. Fixed: endpoint now accepts UploadFile. CSV parsed correctly including SMILES with commas.
GET /api/v1/library/stats and /reagents
Returns real DB stats: 1,216 lipids, 4,600+ activities. Reagent inventory with type filter. Previously 404.
Thread create, message send, markdown response
Requires API key configured in Settings. Thread persistence in SQLite. Suggested prompts functional.
Changelog (9 changes)
- Added pipeline/api/library_routes.py — GET /api/v1/library/stats and /reagents endpoints
- Fixed pipeline/api/iteration_routes.py — /inject-data now accepts multipart/form-data UploadFile
- Fixed web/src/app/results/page.tsx — proper CSV parser for quoted fields (SMILES with commas)
- Fixed web/src/components/sidebar.tsx — removed hardcoded fallback user, proper auth state display
- Fixed web/src/app/formulation/page.tsx — validation preventing submission when percentages != 100%
- Fixed web/src/app/architecture/page.tsx — TypeScript JSX union type error
- Added CI/CD: GitHub Actions workflow with IAP tunnel, docker-compose.prod.yml
- Added python-multipart dependency for FastAPI Form() handlers
- Unignored web/src/lib/ for CI build to resolve @/lib/api imports
Next.js Frontend & TransMA Training Complete
Migrated frontend from Streamlit to Next.js 15 + React 19 with full-featured UI. Completed TransMA model training for lipid property prediction. Built 15 pages covering the complete ICL discovery workflow from pipeline submission to formulation design.
Features
Next.js 15 Frontend
Complete migration from Streamlit to Next.js 15 + React 19 + Tailwind CSS 4. 15 pages with consistent design system, sidebar navigation, JWT authentication, and responsive layout.
TransMA Model Training
3D Transformer + Mamba architecture trained for lipid property prediction. Ensemble of 5 models for uncertainty quantification. CPU-only inference strategy for 16GB RAM.
AI Agent Chat
Claude-powered chat interface with thread persistence, suggested prompts, and markdown rendering. Integrated at /agent with animated gradient border effect.
Architecture Documentation
Comprehensive v3 architecture document served via API endpoint. TOC sidebar with intersection observer for scroll-tracking, collapsible sections, anchor navigation.
Formulation DOE Module
LNP formulation design-of-experiments with component percentage sliders, validation, prediction, and optimization suggestions.
mRNA Design Module
Protein-to-FASTA workflow with tissue selector, modification options, poly(A) configuration, quality metrics, and downloadable FASTA output.
Changelog (7 changes)
- Built web/ directory: Next.js 15 + React 19 + Tailwind CSS 4 + TypeScript
- Created 15 page components: Overview, Pipeline, Runs, Results, Iterations, Chemical Space, Library, Validator, Formulation, mRNA, Conjugation, Agent, Architecture, Settings, Login
- Built pipeline/api/ with FastAPI: 37 endpoints across 12 route modules
- Completed TransMA training — 5-model ensemble for uncertainty quantification
- Added lipid_db/lipid_master.db with 1,216 lipids and ~4,600 activities
- Created architecture_overview_v3.md (76KB comprehensive platform documentation)
- Configured Next.js rewrite proxy for API calls in development
Platform Foundation & ML Pipeline
Initial scaffold of the ICL Discovery Platform. 5-stage ML pipeline (SISSO → Generate → Screen → Conformer → Score) with LSTM generation, AGILE GNN screening, RDKit conformer generation, and TransMA scoring. FastAPI backend with SQLite database.
Features
5-Stage ML Pipeline
DAG-based pipeline: SISSO feature extraction → LSTM molecular generation → AGILE GNN screening (top 500) → RDKit ETKDG conformer generation → TransMA scoring (top 20).
External Model Integration
Integrated 4 external models: SISSO (Fortran, symbolic regression), AGILE (PyTorch GNN, 60k compound pretrained), TransMA (3D Transformer + Mamba), REINVENT4 (generative). Bootstrap script for automated setup.
FastAPI Backend
REST API with JWT authentication, run management (CRUD + progress + cancel/resume), exports (CSV/SDF), and system status monitoring.
SQLite Database
WAL-mode SQLite at lipid_db/lipid_master.db. Tables: lipids, activities, sources, conformations, iterations, iteration_results, users, settings, agent_plans, agent_threads.
Conda Environment
Reproducible icl-discovery conda env: Python 3.11, PyTorch CPU, PyTorch Geometric, RDKit, OpenBabel, gfortran, openmpi. Single environment.yml for all dependencies.
Changelog (9 changes)
- Initial project scaffold with pipeline/, web/, lipid_db/, external/ structure
- Created pipeline_config.yaml — 5-stage DAG definition with model configs
- Built pipeline stages: sisso_stage, generate_stage, screen_stage, conformer_stage, score_stage
- Integrated SISSO, AGILE, TransMA, REINVENT4 as git submodules/externals
- Created FastAPI backend (pipeline/api/main.py) with auth, runs, exports, status
- Set up SQLite database with comprehensive schema for lipid data
- Created environment.yml and setup_env.sh for reproducible environment setup
- Added Dockerfile and docker-compose.yml for containerized deployment
- Created Makefile with dev, api, web, stop, restart targets