π οΈ AI Systems Builder & Quantitative Researcher Β· University of Oxford (Laidlaw Scholar) Β· N1 AI Scholar
π¬ Core Focus: AI Agents & Execution Environments Β· Quantitative Trading & Event Engines Β· LLM Training Dynamics Β· Econometric Data Pipelines
Python Β· PyTorch (Apple Silicon / MPS) Β· Transformer Mechanics Β· Controlled Empirical Study Β· Preregistered Recovery
- Research Problem: Does the temporal order of context lengths presented during training leave a persistent path-dependent deficit on an LLM's final capability, or is the apparent difference an artifact of recency bias and confounded evaluation?
-
What I Built & Key Findings:
- Engineered an inspectable 4.8M-parameter character-level Transformer optimized for Apple Silicon (MPS backend) with modular data pipelines, diagnostic probes, and evaluation harnesses.
- Executed a three-iteration controlled empirical study: evolved from an initial single-seed pilot to a 5-paired-seed design, culminating in a formal preregistered recovery study holding token budgets (4,096 tokens/step), paired seed initializations, data manifests, and AdamW optimizer trajectories invariant.
- Uncovered that an initial severe descending deficit (
$+0.7570$ BPC gap at$T=256$ ) completely reversed to$-0.0485$ BPC after a common$T=256$ recovery phaseβruling out strong persistent path dependence and demonstrating that final performance is dominated by recent context exposure.
- π Repository | π Concise Technical Report (PDF) | π Frozen Preregistration
Python Β· OpenAI Responses API Β· Missing Data Econometrics Β· Multi-API Data Pipeline Β· Folium / Leaflet
- Research Problem: Regulatory invisibility and missing data across micro-enterprises make evaluating social economy density and policy shocks challenging.
-
What I Built:
- Standardized 3,096 master entities across 7 REST/Bulk APIs (Companies House, Charity Commission, FCA, 360Giving, Contracts Finder, IMD) using postcode-blocked Jaccard n-gram matching (
$\ge 85%$ ). - Evaluated OpenAI Responses API (
web_searchtool) vs Chat Completions (gpt-4o-mini), demonstrating a 96% vs 32% accuracy jump and zero-hallucination web extraction. - Modeled the 2013 CIO Policy Shock (+15,600% surge in CIOs, 61% drop in CLGs) and missing data theory (MCAR vs MAR/MNAR).
- Standardized 3,096 master entities across 7 REST/Bulk APIs (Companies House, Charity Commission, FCA, 360Giving, Contracts Finder, IMD) using postcode-blocked Jaccard n-gram matching (
- π Repository | π Live Interactive Maps | π Empirical Paper
Python Β· Interactive Brokers (IBKR) Β· SEC EDGAR API Β· Streamlit Β· Probabilistic Valuation Β· Ongoing / WIP
- Research Problem: Corporate events (earnings releases, guidance updates, Investor Days) shift fundamental cash flows faster than equity markets price them, but naive sentiment approaches suffer from severe look-ahead bias and ungrounded hallucinations.
-
What I Built & Key Edge:
- Implemented an end-to-end quantitative trading engine calculating the Fundamental Revision (
$FR$ ) vs Price Reaction ($PR$ ) Gap to capture structural underreactions across 10-minute to multi-day horizons. - Zero Look-Ahead Bias Ingestion: Pre-event snapshot validation strictly locking consensus expectations prior to event execution; live SEC EDGAR REST API reader with sha256 content hashing and primary citation tracking.
-
Quantified Probabilistic Valuation: Structured thesis engine computing probabilistic Bull/Base/Bear scenarios, Expected Value (
$EV$ ), return spreads, and explicit thesis break conditions. - Automated Execution & Risk Limits: Interactive Brokers (IBKR) paper order routing protected by strict portfolio guardrails (12.5% single-stock cap, 50% portfolio cap, spread <1.5%) and human-in-the-loop Streamlit UI.
- Implemented an end-to-end quantitative trading engine calculating the Fundamental Revision (
- π Repository | π Architecture Specification
Python Β· AST Import Parsers Β· Invariant Baselines Β· Agentic Tool Execution Β· Safe Migration Pipeline
- Engineering Problem: AI coding agents frequently propose hallucinated structural changes, break module imports, and generate unsubstantiated novelty claims without verifiable environment feedback.
- What I Built:
- Developed an enterprise-grade AI Agent Skill and deterministic verification environment for autonomous repository audits and high-conversion landing page restructuring.
- 6-Domain Invariant Verification Baseline (
scripts/invariant_checker.py): Programmatically captures and validates AST module imports, relative Markdown links, package entry points, and CI workflows. - AST Safe Migration Pipeline (
scripts/safe_migrate.py): Performs dry-run simulations, atomicgit mvoperations, and instant automatedgit resetrollback on test or invariant failure. - 3-Layer External Novelty Audit: Orchestrates GitHub Search API, OpenAlex/arXiv API, and web search to output deterministic 6-part proof tuples (
Claim -> Comparable -> Similarity -> Difference -> Evidence -> Confidence).
- π Repository | π¦ Skill Specification
- β‘ gemini-live-multimodal-tutor: Real-time multimodal conversational tutor built with the Gemini Live API for voice/visual problem solving (Google Cloud Hackathon 2026).
- πΊοΈ oxfordshire-population-map: Interactive geospatial portal mapping census demographic distributions and social enterprise density across UK postcodes (
OX1βOX49) (Live Map). - π CrystalNotes: AI transcript intelligence pipeline transforming unstructured audio into deeply hierarchical, structured study notes.
- π tech-launch-promoter-skill: AI Agent Skill for generating developer launch threads, Show HN submissions, and 15-second screen demo scripts.
- π Datawhale_AI4S: Numba JIT-accelerated Kesten stochastic dynamics simulator & active learning candidate phase transition scanner.
- AI Agents & Verification Environments: Deterministic agent execution harnesses, AST-based import/dependency validation, 6-domain invariant checking, tool call schema verification, atomic migration pipelines.
-
Quantitative Trading & Financial Engineering: Event-driven mispricing signals (
$FR - PR$ ), SEC EDGAR live ingestion, abnormal return modeling, probabilistic scenario valuation ($EV$ ), Interactive Brokers (IBKR) API integration, risk management guardrails. - LLM Training Dynamics & Mechanics: Context-length curricula, paired-seed experimental controls, preregistration protocols, BPC validation surfaces, Apple Silicon MPS profiling.
- Quantitative Econometrics & Data Pipelines: Multi-source REST/Bulk ETL pipelines, missing data theory (MCAR/MAR/MNAR), entity resolution (postcode-blocked Jaccard n-gram matching).
- Languages & Frameworks: Python (PyTorch, Numba, NumPy, SciPy, Pandas, Streamlit), JavaScript / TypeScript, HTML/CSS, SQL.
- π Affiliations: University of Oxford (Laidlaw Scholar) Β· N1 AI Scholar
- π§ nanoGPT Research: angelazu-builder/nanoGPT
- π OSEP Research Portal: osep-quant-ai-social-impact
- π Quant Trading Engine (Ongoing): event-driven-trading
- π€ Agent Environments: repo-organizer-skill
- π Interactive Maps: oxfordshire-population-map