Official website for Dr. TMA Pai Endowment Chair - ITIS.
This website showcases our open documentation in the field of Information Technology and Intelligent Systems.
We are committed to advancing knowledge and innovation in technology and education.
Models & Training22
Inference & Deployment13
Serving & Runtime7
Hardware & Systems7
Ecosystems & Tooling6
Theory & Mathematics6
The Unit Distance Problem: What a Model Actually Proved
Theory & MathematicsRamsey Numbers and Extremal Graphs: Two Combinatorics Results
Theory & MathematicsCircuits, Games, and Lattices: Three Complexity Results
Theory & MathematicsA Non-Sofic Group and the Fall of Connes’s Rigidity Conjecture
Theory & MathematicsThree Bounds in the Geometry of Numbers
Theory & MathematicsTen Advances: What OpenAI's Math Results Do and Do Not Show
Theory & MathematicsTokenizers and the Vocabulary Budget
Models & TrainingDesigning Models for Low Precision: Outliers, Scales, and Native FP8
Models & TrainingHybrid Architectures: State Space Models Where Attention Is Too Expensive
Models & TrainingMHA, MQA, GQA, MLA: Attention as a Memory Budget Decision
Models & TrainingInference Engine Architectures: What vLLM, SGLang, TensorRT-LLM, and llama.cpp Optimize
Serving & RuntimeJ-Space: A Verbalizable Workspace Inside Language Models
Models & TrainingWhat Determines Which Language Model You Can Run on Your Device
Inference & DeploymentGroq and Cerebras Accelerators: Deterministic Streaming and Wafer-Scale Compute
Hardware & SystemsInference SLOs: TTFT, TPOT, and Why Throughput Is the Wrong Target
Serving & RuntimeTest-Time Compute: Search, Verification, and Stopping
Inference & DeploymentContinuous Batching: How the Scheduler Decides Your Throughput
Serving & RuntimeMulti-Token Prediction: Training Draft Heads for Faster Decoding
Models & TrainingThe Model Hub as Infrastructure: Transfer, Dedup, and Provenance
Serving & RuntimeScaling Laws: What the Curves Predict and What They Hide
Models & TrainingServing Thousands of LoRA Adapters from One Base Model
Serving & RuntimeModel Compression Beyond Quantization: Pruning, Distillation, and Low-Rank Methods
Models & TrainingModel Alignment Is a Control Problem
Models & TrainingCold Starts: Getting Model Weights from Storage into HBM
Serving & RuntimeCheckpoint Formats: Safetensors, GGUF, ONNX, and the Pickle Problem
Serving & RuntimeLLM Training Optimizers: AdamW, Adafactor, Shampoo, and Muon
Models & TrainingFrontier Mechanistic Interpretability: From Features to Causal Audits
Models & TrainingAI Cluster Networking: NVLink, Infinity Fabric, InfiniBand, and RoCE
Hardware & SystemsDiffusion and Flow Matching: Architectures for Image and Video Generation
Models & TrainingAI Safety Systems in Production: Policy, Permissions, and Validation
Inference & DeploymentLong-Context Models: Position Encoding, Memory, and Retrieval
Models & TrainingSynthetic Data Pipelines: Generation, Filtering, and Validation
Models & TrainingLLM Evaluation: Contamination, Judge Bias, and Production Metrics
Models & TrainingOn-Device Multimodal AI: Text, Vision, and Audio Under a Memory Budget
Inference & DeploymentRetrieval Systems Beyond Basic RAG: Hybrid Search, Reranking, and Evaluation
Inference & DeploymentDistributed LLM Training: Data, Tensor, Pipeline, and Expert Parallelism
Models & TrainingModern Attention Kernels: Tiling, FlashAttention, and Paged Attention
Inference & DeploymentSpeculative Decoding: Draft, Verify, and Accept
Inference & DeploymentModel Context Protocol from First Principles: Messages, Capabilities, and Authorization
Ecosystems & ToolingStructured Generation: How LLMs Produce Valid JSON
Inference & DeploymentHow torch.compile Turns PyTorch Programs into Fused Kernels
Ecosystems & ToolingDisaggregated LLM Serving: Separating Prefill from Decode
Inference & DeploymentInside the KV Cache: Memory Systems for LLM Inference
Inference & DeploymentThe Diversification of AI Hardware: Beyond NVIDIA's Dominance
Hardware & SystemsAMD ROCm for Machine Learning: The Alternative GPU Ecosystem
Hardware & SystemsApple's Metal Ecosystem for Machine Learning: Beyond MLX
Hardware & SystemsApple MLX: A Complete Guide to Machine Learning on Apple Silicon
Ecosystems & ToolingHow Coding Agents Work: Inside the Agentic Loop
Ecosystems & ToolingEdge AI Deployment: Running Models on Constrained Devices
Inference & DeploymentMixture of Experts: Sparse Computation for Efficient LLMs
Models & TrainingRLHF and Preference Tuning: Aligning LLMs with Human Values
Models & TrainingSmall Language Models: When Less is More
Models & TrainingNVIDIA TensorRT and Triton: Production LLM Inference
Inference & DeploymentVoice and Audio AI Models: Architecture, Training, and Deployment
Models & TrainingThe Complete Hugging Face Ecosystem Guide
Ecosystems & ToolingAdvanced Techniques for Local AI Inferencing
Inference & DeploymentFine-Tuning LLMs: A Practical Guide to Tools, Techniques, and Best Practices
Models & TrainingVision Language Models: Architecture, Inference, and Practical Deployment
Models & TrainingUnderstanding GPUs and Parallel Computing
Hardware & SystemsModern Generative AI Guide
Ecosystems & ToolingHPC Guide
Hardware & Systems