Full Table Scan — one open source data tool per issue, installed and run properly.

-- full_table_scan · issue 004

_ □ ✕

Hello, and welcome back to Full Table Scan.

In this edition

- the deep dive: Kubeflow just graduated the CNCF. Tom on what it actually is, what changed this year, and who should run it

- follow the money: Amazon bought the people who write DuckDB, not DuckDB. The $0 team and the $100M tenant

- coming soon from Saiku: Ossie, our agentic layer. Point an agent at your governed cube and it builds the query, runs it, and writes the report

features · vendor: CONCEPT_TO_CLOUD

_ □ ✕

-- featured this issue: Concept to Cloud.

Your product has a blind spot, and you are too close to see it. The Product Blindspot Review from Concept to Cloud puts three senior product and engineering operators on your product for one hour, then hands you a four page written read within five working days: what is working, the three to five blind spots we found, and a prioritised list for the next 30 days. It is a product review, not a code audit, and the session carries no sales pitch. Normally $1,200, and free right now for the launch cohort in exchange for one candid testimonial. If users sign up and drift, or your roadmap is just a pile of competing opinions, apply for a slot.

SELECT body FROM issues WHERE id = 04

_ □ ✕

✓ Showing rows 0 - 0 (1 total, query took 0.0004 sec)

Kubeflow logo

I've worked with Kubeflow since pretty much its inception, while working with Canonical many years ago, and just the other day it graduated to a top-level CNCF project. For those who don't know, the CNCF is the Cloud Native Computing Foundation, and it stewards open-source cloud projects through to a position where they're able to be deployed to the mass market. Other CNCF projects include Kubernetes and Prometheus, to name but a couple, and Kubeflow now also joins the top tier of graduated projects. Kubeflow itself is an orchestration platform that runs inside Kubernetes and turns various machine learning facets into operators within the Kubernetes platform. While Kubeflow isn't for everyone, it does allow machine learning methods of operation to be deployed into organizations that leverage Kubernetes at a scale that is very flexible for the processor on hand. So, if you'd like to learn a bit more, follow along.

What Kubeflow actually is

For anyone new to Kubeflow, it is important to know that Kubeflow is not a single thing. It is an umbrella project that runs a number of different components within the Kubernetes ecosystem that allow machine learning pipelines to be orchestrated. In no particular order, these are:

  • KFP, also known as Pipelines, which is a DAG orchestration layer for ML workflows. This is the component that most people actually mean when they say Kubeflow.

  • Notebooks, which are Jupyter-style workspaces on the cluster. This is still an alpha project.

  • Trainer, which is distributed training, formerly the training operator, but now covers PyTorch, XGBoost, and JAX and Flux.

  • Katib, which is hyperparameter tuning and AutoML.

  • KServe, model serving, which is now actually its own CNCF project.

  • Spark operator, which is Spark on Kubernetes, which is obviously the data processing side of an ML pipeline.

  • Kubeflow Hub, which is a model registry and catalogue.

  • Feast / Kueue, which is a feature store and job queuing, adjacent to a graduation announcement.

There are some dependencies that come with it as part of the deployment, which is actually one of the ongoing niggles with the deployment of Kubeflow. These are Istio, Dex, Cert Manager, and Knative, and that is, as I like to call it, the platform tax.

The mental model

So the way you have to think of Kubeflow is that you're not installing a singular MLOps product. You're extending the Kubernetes API with machine learning nouns. A training job becomes a resource the cluster understands. A pipeline is a CRD. Serving is a CRD.

Now, this is both good and bad. From a positive perspective, everything that you've already got for Kubernetes (assuming that you already have a Kubernetes cluster set up—and I assume you do if you're going into Kubeflow) already applies to your ML workloads. Of course, you get it for free if you use a CNCF project like Argo to deal with the deployment; then you just hook your stuff into it. If you're using RBAC and EKS, you can hook your stuff into it.

I do a lot of work in regulated environments, and this argument is a strong one for allowing data scientists to run their scalable workloads inside an already constrained compute cluster. Similar arguments can be made for healthcare, other fintech environments, and universities and other research institutes that have large-scale clusters for research compute.

Of course, not everything is a plus side, and the abstraction is Kubernetes-shaped. Your scientists who are used to PyTorch and TensorFlow and those types of things now hit Kubernetes-shaped problems, which they will not have seen before. Things like pod scheduling, image pulls, GPU node selectors, and all that type of stuff are new to a data scientist, so the support needs to be there to serve their requirements (to ensure that they're not blocked by weird Kubeflow descriptors).

What changed in the last year

For anyone following Kubeflow—or who's dived in and out of it sporadically over the last however many years it's been running—what's changed in the last year?

The versioning has changed. For anyone who was following the 1.0 line that ended at 1.10, it's now calendar-versioned. For example, 26.03.1 was released on 11 April 2026. What's in that bundle is:

  • KFP 2.16.1

  • KServe 0.18.0

  • Trainer 2.2.0

  • Katib 0.19.0

  • Spark Operator 2.5.0

Running on:

  • Kubernetes 1.36

  • Istio 1.30.1

  • Knative 1.22.0

  • Dex 2.45.1

The GenAI pivot, though, is definitely real, with KServe adding an LLM inference service CRD—so the model registry became Kubeflow Hub, along with the model catalogue and MCP integration.

The support window is roughly six months per community release. If you're deploying this into an operating environment that needs to be able to track these types of things, you need to be able to cope with the support window—obviously, not just from a "how does this thing work" perspective, but also from a security perspective as well.

Where it hurts

So, not everything about Kubeflow is great. Certainly, if you've not worked in a Kubernetes environment before, I would suggest that you probably need to look elsewhere—unless, of course, you really need to run Kubeflow. For those of you who do run Kubernetes environments and want to be able to do machine learning-type science on them, this is a great platform to allow it.

There are, though, some issues with it. For example, there is some community feedback ongoing on the Kubeflow complexity. The complaint is that Kubeflow bundles and tightly couples Istio, Dex, and Cert Manager, which, in practice, pushes you towards a dedicated cluster because, in the cloud-supported ones, they come with it. Users want to be able to swap dependencies and use an operator pattern instead. Maintainers pointed at deployKF, a community distribution, which lets you turn the bundled pieces off and bring your own, which hopefully placates the community in that regard.

In the 2026 Kubeflow SDK user survey, there were four pain points that were uncovered:

  • Infrastructure complexity around GPU scheduling, which has been the bane of Kubernetes since the dawn of time, I think.

  • Debugging, which is scattered logs and cryptic failures, which, again, if you're new to Kubernetes, is something that you have to be able to understand and deal with.

  • The rebuild, push, run loop, killing iteration speed. For those of you who haven't sat in a Docker build chain forever, you won't yet know what it's like to have to wait for everything to build and rerun.

  • Gaps for sophisticated pipelines.

The top roadmap ask was local testing, which, again, for any Kubernetes deployment, is pretty much the same, and the ability to run something before it touches a cluster, because that would greatly speed up the way that this works.

Of course, none of these are complaints of a dying project. They're complaints of a project people are stuck with because it's the only one that does the job, and so they would like to see it improved. This is where stewardship with the CNCF really comes to bear fruit.

Our opinion

So, there are obviously pros and cons to using Kubeflow. As we have pointed out, there are a lot of complexities with just running Kubernetes in general, doing it properly, and following best practices (to ensure that you don't end up with just a bunch of junk running in your cluster).

If somebody wants pipelines, there's a reasonable chance that you deploy this thing and then, six months later, there's an Istio mesh nobody chose, a Dex config that nobody understands, and a cert manager upgrade is blocking the quarter.

The dividing line, though, isn't machine learning maturity. It's whether you already run a Kubernetes cluster in production with a platform team that owns it. If you do, Kubeflow is close to free because you're paying for Kubernetes anyway. You've got all the best practices in place. You've got people that understand how that cluster works, and so you can deploy that thing into a scalable environment and run machine learning experiments at huge scale, not just what's on a developer's laptop. If you don't, then you've bought a second full-time job, and it's called MLOps.

We deployed this into a fintech provider a year or so ago, and it provided a great amount of scale and experimentation improvements that really drove that organisation forward because it allowed them to run experiments at scale in a flexible environment. This was without all the other hookups that needed to happen to allow the platform team to run it, as they already had Kubernetes in place.

The verdict

If you're interested in experimenting with Kubeflow, one suggestion is to adopt Kubeflow's components but not the platform. Take Pipelines, KServe, or Trainer on their own terms. They're good, and now they're CNCF-graduated and safe to build against, because you know they will continue to be maintained for the foreseeable future.

Take the full bundled distribution only if you already have a platform team that would have run Istio anyway. Everyone else should run a managed distribution or a hosted alternative, and revisit when the local testing and modularity work lands.

SELECT * FROM the_money_table

_ □ ✕

✓ a brief from Juan Resendiz

DuckDB logo

Amazon bought DuckDB last week. Except it didn't.

It bought DuckLabs, the thirty or so people in Amsterdam who actually write DuckDB. The database itself never changed hands. It is still MIT licensed, still owned by a Dutch nonprofit called the DuckDB Foundation, still free this morning. That gap, between the code and the people who make it, is the whole money story.

Here is the part that gets me. DuckLabs never took a cent of venture capital. Five years, bootstrapped, owned by its founders and its staff, paid for out of support contracts. The VCs were calling early and they said no. They were turning down buyers back in 2022, before most of us had heard the word DuckDB.

Now look at the company built on top of them. MotherDuck raised about $100M and last carried a $400M valuation, selling DuckDB as a service. So the people who build the engine took nothing, and the people who rent it out took nine figures. One day before the Amazon news, MotherDuck went and bought a startup of its own. The day after, it announced it will now sell DuckDB support directly, competing with whatever AWS ships next, running on AWS. Everyone is renting a nest they do not own.

Why does Amazon want the authors badly enough to hire all thirty? Volume. DuckDB does more than a million downloads a day and past 50 million a month on PyPI, and it reads a Parquet file or an S3 bucket straight, with no warehouse in the middle. If S3 becomes the place you query and not just the place you store, you want to own the thing doing the querying. Amazon had already pushed 2.5 billion queries through its DuckDB integration inside Quick before it ever bought the team.

The price is undisclosed and I am not going to invent one. But we have watched this exact trade three times now. Databricks paid roughly a billion for Tabular, then roughly a billion again for Neon. Snowflake paid about $250M for Crunchy Data. Every single time the code was already open, so the thing you are buying is the maintainers. Amazon will not say the number. The comps tell you the neighborhood.

My take: you cannot rent a foundation. DuckDB stays open because a nonprofit holds the IP, and Amazon has every reason to keep it that way, since a bigger DuckDB just means more compute sitting on S3. The bet that should keep you up at night is not the license. It is that the roadmap now lives inside the same company that sells you the storage underneath it.

This is also why we keep DuckDB close at Saiku. It reads a file or an S3 bucket without a warehouse tax, and you can put a live pivot straight on top of it or hand the same governed model to an agent. Own your engine, keep it portable, and it stops mattering who buys whom.

COMING SOON · vendor: SAIKU

_ □ ✕

Next from Saiku: point an agent at your cube.

Saiku 4.8 is in the release candidates now, and the headline is Ossie, our agentic layer. You give it a question, it builds the query, runs it against your governed semantic model, and writes the report back. It is the same cube your Excel pivots and dashboards already read, now reachable over MCP, so any AI client asks the model and gets a governed answer instead of a guess. We even shipped sample agents in Python and TypeScript so you can wire your own in an afternoon. If the DuckDB story above is about who owns the engine, this is the other half: own the semantic layer your agents read, and it stops mattering whose model is asking. The sample agents are on GitHub.

Tom Barber
Tom Barber signature

Tom Barber

// Concept to Cloud · Saiku @ Spicule

Juan Resendiz
Juan Resendiz signature

Juan Resendiz

// Concept to Cloud · works on Saiku

INSERT

_ □ ✕

One open source data tool per issue, installed and run properly. Forward this to whoever is about to rebuild a semantic layer by hand.

1 row affected. Unsubscribe whenever.