groupby on a single column in Python Pandas returns a series. Want a single-column data frame? Instead of pass a one-element list of data columns:
Invoke “groupby” with two categorical columns in Python Pandas, and get a two-part multi-index: Turn into a data frame with unstack:
A “groupby” call in Python Pandas is normally sorted by index. But if you’re grouping by month name, April will come before January. Pass “sort=False”, and the index will reflect the order of appearance:
The simplest grouping in Python Pandas is groupby: For example: Returns a series whose index is the unique values from passenger_count.
You can remove NaN from a Python Pandas data frame with dropna, but be careful: It removes rows with even one NaN, which can be overkill. Pass “thresh” to allow some rows to remain, even if they contain NaN:
Coming to Python Pandas from NumPy? You’ll reach for np.isnan: Unfortunately, this works. Better, use s.isna (or s.isnull). But the best way to drop NaN? Use dropna:
PyArrow dtypes in Python Pandas are nullable (with pd.NA): s is: 0 101 <NA>2 30dtype: int64[pyarrow] The dtype is int64, but allows nulls. (Use np.nan? It’s turned into pd.NA.)
You have a Python Pandas series with ints + NaN. You don’t want float forced on you. Solution: Use the “extension” type Int64 (note Initial Caps) and pd.NA: s is: 0 101 <NA>2 30dtype: Int64
Some of the most satisfying work I do happens one on one. You arrive with a real problem from your real job — code that will not behave, an architecture you are unsure about, a Git situation that has you stuck — and an hour later it is smaller, or gone. I have offered these…
Missing data? NumPy calls it nan. Python Pandas displays it as NaN. But: Pandas doesn’t define pd.nan or pd.NaN. NumPy removed np.NaN in version 2.0. So you have to refer to np.nan from within Pandas, and it’ll be displayed as NaN.