Computer Science and Engineering, University of Michigan

OOPSLA 2026

Interactive Data Analysis
with Lively Typed Tables

Alexander Bandukwala · Cyrus Omar

University of Michigan's Future of Programming Lab

Exploratory data analysis today

Exploring, cleaning and reshaping data in a live programming environment: writing code and running it are interleaved.

In [1]:
import pandas as pd df = pd.read_csv("weather.csv") wide = df.pivot(index="date", columns="element", values="value")
In [2]:
wide["wide["tmean"]
Out:
Traceback (most recent call last) File pandas/core/frame.py:3805, in DataFrame.__getitem__(self, key) KeyError: 'tmean'

but that feedback is

  • mostly autocomplete and suggestions
    few diagnostics or error marks on the code
  • based on the current runtime state
    from the values the kernel holds, not the code
  • only there for global variables
    limited support inside of functions
  • dependent on cell execution order
    cells run or edited out of order give non-reproducible feedback for analyses†

† Out-of-order execution in notebooks: Pimentel et al., MSR 2019 · Chattopadhyay et al., CHI 2020

Types for table schemas

In [1]:
df = pd.read_csv("weather.csv") df
Out[1]:
dateelementvalue
"2010-01-30""tmax"27.8
"2010-01-30""tmin"14.5
"2010-02-02""tmax"27.3
"2010-02-02""tmin"14.4
In [2]:
wide = df.pivot(index="date", columns="element", values="value") wide
Out[2]:
datetmaxtmin
"2010-01-30"27.814.5
"2010-02-02"27.314.4
In [3]:
wide["tmean"]
Out[3]:
KeyError: 'tmean'
1 What columns does wide have?

The column values are being promoted to the schema.

2 Can "tmean" be rejected statically?

Fancy type systems

One answer is to make the type system richer

Type system features

  • Record operations, row polymorphism Cardelli & Mitchell 1991 · Wand 1987 · Rémy 1994
  • Heterogeneous lists, dependent records Kiselyov et al. 2004 · Waszczuk et al. 2017
  • Match types, named tuples Blanvillain et al. 2022 · Odersky 2024

Applied to tables

  • Dependently typed tables Lean: Rotella 2024 · Idris: Wright et al. 2022
  • Typed Spark datasets Scala: Frameless (Typelevel)

These are powerful: they can type the operations tables need.

But value-dependent operations need a full scan of the dataset.

And in exploratory analysis, you don't yet know what's in that data.

Our approach

  1. Use static typing when convenient
  2. Fall back to gradual typing when our static types are not enough buys expressivity · costs feedback
  3. Introduce live typing provides the feedback static types cannot
the Hazel editor

We build on Hazel, a live and gradually typed programming environment.

Hazel Lab Hazel + structural operations + live typing + rich probes

Tables as lists of labeled tuples

Brown Benchmark for Table Types

B2T2 Lu, Greenman & Krishnamurthi 2022

Example tables, programs and error scenarios — and a Table API of 49 operations.

Can our type system express the operation?

Can it enforce its constraints?

one operation from the Table API

tsort : t1:Table × c:ColName × b:Bool → t2:Table
requires
c ∈ header(t1)
ensures
schema(t2) = schema(t1)
nrows(t2)  = nrows(t1)

B2T2: Hazel

operations expressible preconditions enforced postconditions enforced
Hazel 17 / 49 4 / 72 29 / 142
  • Labels are not first classA label cannot be passed to a function, returned from one, or computed.
  • No polymorphism over schemasNo static type that accepts tables with different schemas.

one operation from the Table API

tsort : t1:Table × c:ColName × b:Bool → t2:Table

Labeled tuple to list conversions

  • to_lvs: tuple to listEach entry becomes a label–value pair, the label a string.
  • from_lvs: list to tupleBuilds a tuple back from label–value pairs.
  • to_lvs returns[(label=String, value=τ)]
    τ is the meet of the entry types
  • from_lvs returns ?Labels are determined at runtime

B2T2: Hazel Lab

operations expressible
of 49
preconditions enforced
of 72
postconditions enforced
of 142
Hazel 17429
Hazel Lab 49*5 38

Every operation is at least partially expressible, but most constraints go unchecked.

The paper also reports an intermediate ablation, Hazel with tuple extension.

* four only partially expressible, e.g. leftJoin requires the right table to have at least one row

Live typing

In a live programming environment, we have dynamic values. We should use them to bolster static types.

  1. Observe the values that reach each statically underspecified expression
  2. Synthesize the type of each observed value
  3. Combine those types with their meet falling back to the static type if they conflict
  4. Re-check the program with the unknowns filled in

Statically, this program has no errors

? is consistent with every type, so static analysis has nothing to report.

Is this just dynamic typing?

The opening example again, with a report that is never called.

  • Checks code that never ranreport is never called.
  • Marks in the editorOn the code, not in a stack trace.

Back to the benchmark

operations
expressible
of 49
preconditions
enforced
of 72
postconditions enforced
of 142
without runningafter running
(live typing)
Hazel 1742929
Hazel Lab 49*538 96†

No extra annotations, and no richer types in the surface language.

Live typing doesn't help preconditions.

The paper also reports an intermediate ablation, Hazel with tuple extension.

* four only partially expressible, e.g. leftJoin requires the right table to have at least one row

† assumes each call returns a table with the same schema every time it runs

Rich probes for tables

A domain-specific view of a probed value, where direct manipulation edits the underlying syntax.

User study

An exploratory lab study with users familiar with statically typed functional programming. Not intended to assess transfer to data scientists without training.

  1. Can users use Hazel Lab's table operations to do basic data transformation and data cleaning tasks?
  2. Can users effectively use rich probes in data cleaning and transformation tasks? What problems do they face?
  3. Do users find live typing errors useful in identifying programming errors, and helpful in understanding program behavior?

Study design

  • Seven participants, 9 to 30 years of programming
  • Two-hour remote sessions
    • Hazel tutorial
    • Six tasks
    • Exit survey
  • Within-subjects design
    • Matched task pairs, off then on
    • Order counterbalanced

Adoption intent questions

I would want…
to use this labeled tuple representation in my own environment μ = 5.4
rich probes available in my programming environment μ = 5.7
a programming environment that shows live types μ = 5.9
a programming environment that shows live type errors μ = 5.9
Strongly disagree (1) Disagree Somewhat disagree Neutral Somewhat agree Agree Strongly agree (7)

Every response at least neutral: 24 of 28 positive, none negative.

All seven were already fluent in statically typed functional programming.

Not intended to assess transfer to data scientists without that training, many of whom work in Python or R.

Related work

live typing

  • Type providers Syme et al. 2012A compiler plugin builds types from the data source.
  • The Gamma Petricek 2017Lazily generates methods for dot-driven development, one pipeline step at a time.
  • RubyDust An et al. 2011Infers static types from test runs, as a batch step before type checking.
  • REPLs, Jupyter Pérez & Granger 2007Feedback is driven by the kernel's current state.
  • Cuis Smalltalk Wilkinson 2019 · Ostrovsky 2024The VM tags each variable with the set of classes it has held. A checker flags incompatible method calls.

rich probes

  • Livelits Omar et al. 2021GUIs that fill a hole with a literal of one fixed type.
  • The Gamma Petricek 2017 · 2022The table preview doubles as an editor, with buttons to add or remove steps in the chain of method calls.
  • Histogram Petricek 2019Direct manipulation of table previews. The program is a log of interactions, and functions keep their sample inputs.
  • Mage Kery et al. 2020A table widget can be summoned on a notebook variable, and its actions are written back into the cell as code.

Future work

live typing

  • Stream results into the typechecker as they are produced, so a mark can appear during a long computation
  • Integration with fancy types
  • User-configurable granularity, instead of on or off for the whole program
  • Other treatments of inconsistent observations, which keep their static type today

rich probes

  • User-definable and composable rich probes, reconciled with livelits
  • Actions for probes on patterns, read-only today

Scalability

  • Incremental typechecking and evaluation, for notebook-like responsiveness
  • Columnar table representations
  • A foreign data interface: data stays in external storage but behaves like a table in the code
  • Move computation to the data, such as by compiling Hazel data pipelines to Apache Spark Apache Foundation 2025
  • Productionalization: performance engineering to move the research prototype to a real tool
Computer Science and Engineering, University of Michigan

summary

Hazel Lab

hazel.org/build/tables-study

Hazel Lab in your browser, start with the tutorial

Artifact on Zenodo: 10.5281/zenodo.21458220

QR code for hazel.org/build/tables-study

Structural operations on labeled tuples

The same operations on a row whose type is unknown: let x : ? = (a=1, b=2, 3, c=4)

exampletypevalue
(a=1, b=2) ... (b=3, c=4)⇒(a=Int, b=Int, c=Int)(a=1, b=3, c=4)
project_labels((a=1, b=2, 3, c=4), `a`, `b`)⇒(Int, Int)(1, 2)
omit_labels((a=1, b=2, 3, c=4), `b`, `c`)⇒(a=Int, Int)(a=1, 3)
omit_all_labels((a=1, b=2, 3, c=4))⇒(Int, Int, Int, Int)(1, 2, 3, 4)
exampletypevalue
x ... (b=3, c=4)⇒?(a=1, b=3, 3, c=4)
project_labels(x, `a`, `b`)⇒(?, ?)(1, 2)
omit_labels(x, `b`, `c`)⇒?(a=1, 3)
omit_all_labels(x)⇒?(1, 2, 3, 4)

These are primitives rather than library functions, because labels are not first-class values in Hazel.

When the argument's type is unknown, so is the result's.

project_labels still knows how many entries it returns, because the labels are written in the source.

The values are unaffected.

Each result is computed correctly at runtime. Only the static type has lost the information.

Tabular programming systems

Scala
+ IDE
PolynoteJupyter
+ pandas
Hazel
Lab
function barrierno type-directed services inside a function●◐○●
distal feedbackerrors reported far from the cause, and only after running◐◐○◐
external schemadata parsed at runtime, so its schema is invisible to static analysis◐◐◐●
expressivitystructural table operations are hard to type statically◐◐◐●
out-of-order executionfeedback from execution history rather than from the program●◐○●

No existing system addresses all five. Our own row is partial on distal feedback: live typing shortens the spatial distance and not the temporal one — you still wait for the program to finish.

Why the meet?

  • Hazel is a static-first gradually typed language. The unknown type exists mainly to give a type to holes, and there is no typecase operation.
  • So we assume programs are converging toward more specific types, and take the greatest lower bound with respect to precision.
  • In a system with subtyping, or a dynamic-first language, the join would be the right choice instead.
(Int, ?) ⊓ (?, String)  =  (Int, String)
(Int, ?) ⊔ (?, String)  =  (?, ?)
?
↓
(?, ?)
↙  ↘
(Int, ?)(?, String)
↘  ↙
(Int, String)

More precise as you go down. The meet of two observations is the highest type below both.

The unknown type is not any

Hazel

let x : ? = "" in
let y = [x, 1] in …

y : [Int]

The unknown type is a placeholder for something more precise, so it gives way to the Int beside it.

TypeScript

const x : any = "";
const y = [x, 1];

y : any[]

any is absorbing: it propagates outward and erases the information its neighbors carried.

How to_lvs is typed

A static error if the argument is not a tuple. Otherwise the value type is the meet of the entry types, falling back to the unknown type when they are inconsistent.

to_lvs((a=1, b=2, c=3))            ⤳ [(label=String, value=Int)]

to_lvs((a=?, b=1, c=?))            ⤳ [(label=String, value=Int)]

to_lvs((a=true, b=1, c=2.0))       ⤳ [(label=String, value=?)]

to_lvs((a=(true,?), b=(?,?), c=(?,1.0)))
                                   ⤳ [(label=String, value=(Bool, Float))]

The last case is the interesting one: the meet is computed pointwise, so partial information from several entries combines into a type more precise than any single entry supplied.

One function, two datasets

pinned to the first invocation
pinned to the second invocation

Applied to two datasets, the schemas are inconsistent, so no meet exists and the result keeps its static type — no marks. Each probe collects one value per invocation, so pinning one invocation restricts live typing to that call: the mark lands on Unknown for the first dataset and moves to Recreational for the second.

What makes this work in Hazel

Column names as data

nameage quiz1quiz2 midterm quiz3quiz4 final average-quiz
"Bob"12897779878.25
"Alice"17688887857.25
"Eve"13798488778.00
buildColumn(gradebook, "average-quiz", (row) => {
  const quizCols = header(row)
    .filter((c) => startsWith(c, "quiz"));

  const scores = quizCols
    .map((c) => getValue(row, c) as number);

  return sum(scores) / length(scores);
});

Add a column of mean quiz scores. Adapted from B2T2.

1 What is the type of gradebook?
2 What about row?
3 Which columns are the quiz columns?
4 What type does getValue(row, c) return?

Statically typing this requires a sophisticated type system.

Live typing on buildColumn

The gradebook example: a new column named by a string, computed from the columns whose names start with quiz.

--:--
→ / space next
← back
1–9 jump to slide
s speaker notes
t start / pause timer   r r reset
[ ] presenter: shrink / grow the panel
esc reclaim focus from the editor
? this help