John Memon

New York, NY

Summary

AI/ML engineer specializing in applied GenAI systems, multi-agent orchestration, and GPU-level performance work. Takes ambiguous problems to shipped products as sole owner, with depth across computer vision, NLP, real-time systems, and model interpretability.

Experience

All In On Data — AI Engineer Jul 2024 – Current
Autolog – AI-Powered Knowledge Acquisition Platform
Client Engagements — Pre-Sales · Technical Advisory and Delivery
Philips — ML Engineering Co-op Jan – Aug 2023
Northeastern University — Research Assistant Sep 2022 – Jan 2023
Morse Corp — Data Science Co-op Jan – Jun 2022

Projects

Abstruct: An AI-enabled UX for memorization and concept engagement that turns ideas into content — GenAI-powered Socratic recall using FSRS scheduling, real-time voice input (Deepgram Nova-2), continuous LLM evaluation, and tree-structured conversation branching; currently designing a handheld companion device (Raspberry Pi Zero 2W, e-ink display, camera OCR, MEMS microphone) for screen-free capture of primary-source material with user commentary
Fuel-Code: Developer-productivity platform built on a deep understanding of Claude Code internals — models the full agent lifecycle (sub-agents, teams, skills, worktrees, permission modes) to preserve session transcripts and let developers reuse and iterate on productive sessions
Arxiver: RAG-LLM research tool that spins up expert chatbots in user-specified domains
CUDA Deep Learning: From-scratch neural network in CUDA and C++ — hand-written kernels for matrix ops, activations, and losses; forward + backward passes on-device; multi-architecture builds via Makefile fat binaries targeting Volta (sm_70) and Pascal (sm_61)
SimpleProcessor: Designed and simulated a custom CPU in Verilog with a C++ test harness — implemented the full instruction fetch/decode/execute pipeline, register file, and ALU, working hands-on across the hardware-software boundary
MemoryAllocator: Custom memory allocator in C — free-list management, sbrk/mmap-backed regions, and fragmentation analysis, validated by a dedicated test harness; performance-critical low-level memory work
NUFileSystem: FUSE-based filesystem in C — block allocation, inode tables, and directory-tree traversal implementing the POSIX interface, with a full test suite covering the syscall surface

Education

Northeastern University — BS Mathematics, Cum Laude, 3.70 GPA
Eta Kappa Nu — EECE Honors Society
Coursework: Machine Learning, Probability, Linear Algebra, Real Analysis, Group Theory, Algorithms, Computer Systems, Theory of Computation
Certification: Google Cloud Professional (Oct 2025)

Skills

Languages: Python, TypeScript, SQL, C/C++, CUDA, Verilog/HDL, Go
AI/ML: PyTorch, PyTorch Lightning, LangChain, Anthropic SDK, MCP, DeepSpeed ZeRO, LoRA/PEFT, OpenCV, scikit-learn, NumPy, Cython, nltk, Deepgram
Systems & Performance: Custom CUDA kernels, multi-architecture GPU builds (Volta/Pascal), manual memory management, FUSE-style filesystems, processor / HDL design, real-time streaming pipelines, low-latency WebSocket / WebRTC protocols
Infrastructure: Docker, Redis, PostgreSQL, AWS (EC2/S3/EBS), Supabase, Cloudflare R2, nginx, Railway, Vercel
Web & Real-Time: React, Hono, Express, Bun, WebSocket, WebRTC, Chrome DevTools Protocol, Server-Sent Events
Observability & Tooling: Prometheus, Loki/Promtail, Pino structured logging, NVTX, torch.cuda.memory profiling, Turborepo, Drizzle ORM, Git, pytest

Interests: Guitar, staying active, time with loved ones.