Reading a CSV into a Python Pandas data frame? Use the “dtype” keyword arg and a dict to specify dtypes, and avoid the int64/float64/str defaults:
Want to change the dtype of a Python Pandas series? You can’t — at least, not by assigning: Instead, use “astype” to get a new series, and replace the old one:
Get the dtype of a Python Pandas series with dtype: Get the dtypes of all columns in a data frame with “dtypes”, which returns a series: df.dtypes How many of each dtype? Use value_counts:
Want to get a quick look at a Python Pandas data frame you just created? Option 1, use df.head(n) to look at the first n rows: Option 2, use df.sample(n) to look at n randomly selected rows:
If you want to get better at Pandas, the hard part isn’t finding tutorials. It’s finding problems worth solving. Most exercises hand you a tidy little table of five rows and ask you to sum a column — which teaches you the syntax, but nothing about the job. For the last 3.5 years, I’ve written…
Want to read a zipped CSV file into a Python Pandas data frame: Just pass the filename to read_csv: The file can contain a single CSV file, with an extension of .gz, .bz2, .zip, .xz, .zst, .tar, .tar.gz, .tar.xz, or .tar.bz2.
What sheets are in an Excel document you’re about to read into Python Pandas? You can check with: You’ll get a list of Python strings back.
Which is faster in Python Pandas, | or isin? I prefer isin; it’s clearer to write and read. Plus, fewer worries about parentheses. But is it faster? Depends on the dtype! isin edges out | on strings, but | wins on ints. Bottom line: Use %%timeit to check. Don’t just guess!
Want rows in a Python Pandas data frame that might have several values? Another way (besides the | I showed yesterday) is the “isin” method: We’ll compare isin vs. | speed tomorrow. But I find this far more readable.
Want rows in a Python Pandas data frame that might have several values? You can use the | operator, but be sure to use () to avoid precedence issues: