What the Data Crowd Was Reading in July 2026
Tools, techniques and deep dives worth reading that I came across in July 2026.
Fellow Data Tinkerers
It’s time for another round-up on all things data and AI!
But before that, I wanted to share with you what you could unlock if you share Data Tinkerer with just 1 more person.
There are 100+ resources to learn all things data (science, engineering, analysis). It includes videos, courses, projects and can be filtered by tech stack (Python, SQL, Spark and etc), skill level (Beginner, Intermediate and so on) provider name or free/paid. So if you know other people who like staying up to date on all things data, please share Data Tinkerer with them!
Without further ado, let’s get to the round up for July!
Data science & AI
Kimi K3 Architecture Notes (4 minute read)
Sebastian Raschka, PhD dissects Kimi K3’s sparse architecture, attention residuals, NoPE design and inference-efficiency trade-offsThe Cheapest Way to Make Open Weight AI Models Better (27 minute read)
Devansh explains why targeted post-training can improve open-weight models more economically than scaling pre-training alone.A Taxonomy of Self-evolving Agents (14 minute read)
Shilong Liu proposes a practical taxonomy for agents that adapt their memory, tools, prompts or policies through experience.The LLM Critics Are Right. I Use LLMs Anyway. (19 minute read)
Jeremy Theocharis accepts the strongest criticisms of language models while explaining the disciplined uses that still make them valuable.Closing the Verification Loop (12 minute read)
Kieran Klaassen argues that agents become dependable only when verification is built into the loop rather than left to a final human check.How an AI Token Travels Through a Data Center (26 minute read)
Chris Zeoli traces one token through gateways, schedulers, KV caches, accelerators and networks to explain the real mechanics and economics of inference.How to Implement a Unified Memory From Scratch (16 minute read)
Paul Iusztin builds a unified memory layer from scratch and explains how agents can retain useful state without turning context into clutter.
Data engineering
I Spent 10 Hours Learning Multithreading and Multiprocessing (14 minute read)
Vu Trinh compares multithreading and multiprocessing through hands-on examples so Python practitioners can choose the right concurrency model.The Data Anarchy Tax: Why your team is firefighting 45% of the time (9 minute read)
Good article by Joe Reis on why fragmented ownership and unclear operating rules create a measurable data-anarchy tax that keeps teams in reactive work.The Most Expensive Mistake in System Design (12 minute read)
Erfan Hesami explains why designing for imagined scale too early can make data systems harder and more expensive to operate than the problem requires.The Startup’s Postgres Survival Guide (11 minute read)
Alexander Belanger lays out pragmatic Postgres choices for startups that need reliability without prematurely adopting a more complex data platform.Your Data Warehouse Isn’t Integrated Just Because the Tables Are in One Place (12 minute read)
SeattleDataGuy argues that co-locating tables is not integration when semantics, contracts and ownership remain fragmented.So, is data modeling dead? (16 minute read)
Daniel Beach examines what AI-assisted development changes about data modelling and why deliberate structures still matter.Backfilling: The Most Underrated Feature of a Data Pipeline (10 minute read)
Lucie Thimus explains why safe, repeatable backfills are a core capability for correcting history and evolving production pipelines.FastAPI for Model Serving: The Standard You Should Actually Use (12 minute read)
Andres Vourakis shows how to structure a FastAPI service for model inference with production concerns such as validation and deployment in mind.
Data analysis and visualisation
The Model Is Smart. Your Company Is the Problem (10 minute read)
Olga Berezovsky argues that failed AI analysis is often rooted in unclear metrics and organisational fragmentation rather than model capability.Going Bayesian Automates Your Manual Data Analysis (9 minute read)
Eric J. Ma shows how robust likelihoods and hierarchical Bayesian models can replace manual outlier handling with a reproducible analytical process.
Other interesting reads
How Tech Workers Are Feeling in 2026: A Workforce Splitting in Two (22 minute read)
Interesting analysis by Noam Segal and Lenny Rachitsky using a large sentiment survey to show how AI is dividing tech workers between amplification and destabilisation while burnout rises.You Don’t Have to Be Smart If You Can Think Clearly (4 minute read)
Sean Goedecke argues that methodical, clear thinking under uncertainty is more dependable than flashes of intuition when engineers face genuinely hard problems.
Quick favor - need your take
Was there any standout article or topic from July I missed? Feel free to drop a comment or hit reply, even a quick line helps.
If you are already subscribed and enjoyed the article, please give it a like and/or share it others, really appreciate it 🙏
Keep learning
What the Data Crowd Was Reading in June 2026
It’s time for another data/AI roundup and here are the highlights from June👇
𝐃𝐚𝐭𝐚 𝐒𝐜𝐢𝐞𝐧𝐜𝐞 & 𝐀𝐈
Why AI systems need different monitoring than web services
How to tell when something is actually a RAG problem
Why stateful swarms may make agents cheaper and smarter
When trees still beat tabular foundation models
Why running local models is finally practical
𝐃𝐚𝐭𝐚 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠
Why vibe coding is risky but agentic engineering can work
How to build a simple, bulletproof data pipeline
Broker-visible vs client-local parallelism
Basic Spark concepts worth knowing first
Why data fundamentals matter more than ever in 2026
𝐃𝐚𝐭𝐚 𝐀𝐧𝐚𝐥𝐲𝐬𝐢𝐬 & 𝐁𝐈
Why data quality depends on the decision it supports
Why general AI models still struggle with data analysis
What the Tableau exodus says about the future of BI
Plus: why AI hurts big consulting more than specialist expertise and practical advice for junior data candidates trying to stand out in a brutal job market.
What the Data Crowd Was Reading in May 2026
It’s time for another data/AI roundup and here are the highlights from May👇
𝐃𝐚𝐭𝐚 𝐒𝐜𝐢𝐞𝐧𝐜𝐞 & 𝐀𝐈
How AI agents improve through feedback and evals
A practical guide to choosing the right graph model
How memory works inside AI agents
How to build and evaluate reliable Claude Skills
What matters when taking RAG into production
𝐃𝐚𝐭𝐚 𝐄𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠
The five very different worlds of data engineering
Why simpler infrastructure beat Databricks
Why ETL versus ELT is mostly the wrong debate
Building a sub-50-cent ETL pipeline on AWS
Giving analytics agents context only when needed
𝐃𝐚𝐭𝐚 𝐀𝐧𝐚𝐥𝐲𝐬𝐢𝐬 & 𝐕𝐢𝐬𝐮𝐚𝐥𝐢𝐬𝐚𝐭𝐢𝐨𝐧
Where AI genuinely helps analysts and where it falls short
Why polished AI dashboards can still be useless
How to show plan-versus-actual gaps clearly
Plus: which jobs AI may cut next, what China’s AI labs look like from the inside and why superstar AI researchers can command enormous salaries.












