{ "cells": [ { "cell_type": "markdown", "id": "plotting-00", "metadata": {}, "source": [ "# Plotting tutorial\n", "\n", "ProteoPy's `pr.pl` functions help you inspect sample coverage, missing\n", "measurements, abundance distributions, and patterns across experimental\n", "groups. This tutorial is an overview: each example asks a small question,\n", "while the [plotting API](https://proteopy.readthedocs.io/en/latest/api/pl.html)\n", "documents the full set of options.\n", "\n", "**General plotting concepts**\n", "\n", "An AnnData object stores samples in rows and proteins or peptides in columns.\n", "`.X` contains measured intensities, `.obs` contains sample annotations, and\n", "`.var` contains feature annotations. Sample labels come from\n", "`.obs[\"sample_id\"]`; protein labels come from `.var[\"protein_id\"]`.\n", "Detection counts describe **coverage**, whereas intensities describe\n", "**abundance**. More detected proteins does not necessarily mean that the\n", "proteins shared between samples are more abundant.\n", "\n", "We use the protein-level erythropoiesis data from\n", "[Karayel et al. (2020)](https://doi.org/10.15252/msb.20209813), covering five\n", "differentiation stages. The loader downloads and caches the data on first\n", "use. \n" ] }, { "cell_type": "code", "execution_count": 1, "id": "plotting-01", "metadata": { "execution": { "iopub.execute_input": "2026-09-24T14:12:23.878029Z", "iopub.status.busy": "2026-09-24T14:12:23.877882Z", "iopub.status.idle": "2026-09-24T14:12:29.065151Z", "shell.execute_reply": "2026-09-24T14:12:29.064782Z" } }, "outputs": [ { "name": "stderr", "output_type": "stream", "text": [ "/opt/homebrew/Caskroom/miniforge/base/envs/proteopy1/lib/python3.11/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html\n", " from .autonotebook import tqdm as notebook_tqdm\n" ] }, { "data": { "text/plain": [ "AnnData object with n_obs × n_vars = 20 × 7758\n", " obs: 'sample_id', 'cell_type', 'replicate'\n", " var: 'protein_id', 'gene_id'" ] }, "execution_count": 1, "metadata": {}, "output_type": "execute_result" } ], "source": [ "import matplotlib.pyplot as plt\n", "import numpy as np\n", "import pandas as pd\n", "import proteopy as pr\n", "\n", "adata = pr.datasets.karayel_2020()\n", "raw_intensities = adata.X.copy() # Check preservation at the end.\n", "adata" ] }, { "cell_type": "markdown", "id": "plotting-02", "metadata": {}, "source": [ "**Ordering samples and groups**\n", "\n", "The intended ProteoPy convention is lexicographic ordering of string-coerced\n", "labels for ordinary string/object annotations (`F1`, `F10`, `F2`), and\n", "category order for categorical annotations. Store a biological sequence\n", "as a pandas categorical rather than relying on the object's row order.\n", "An explicit `order` takes precedence over metric sorting such as\n", "`ascending`; otherwise the annotation order applies. Dendrograms impose\n", "their own order when clustering is enabled.\n", "\n", "
Current ordering exceptions
\n", "Some functions still retain input order or sort by counts. In particular,\n",
"n_proteins_per_sample can retain row order within groups, and\n",
"n_samples_per_category sorts by counts by default. Their current\n",
"order handling can append unlisted items rather than subset\n",
"them. These examples pass complete explicit orders where needed; consult\n",
"each function's API before relying on subsetting.
| \n", " | sample_id | \n", "cell_type | \n", "replicate | \n", "
|---|---|---|---|
| LBaso_rep1 | \n", "LBaso_rep1 | \n", "LBaso | \n", "rep1 | \n", "
| LBaso_rep2 | \n", "LBaso_rep2 | \n", "LBaso | \n", "rep2 | \n", "
| LBaso_rep3 | \n", "LBaso_rep3 | \n", "LBaso | \n", "rep3 | \n", "
| LBaso_rep4 | \n", "LBaso_rep4 | \n", "LBaso | \n", "rep4 | \n", "
| Ortho_rep1 | \n", "Ortho_rep1 | \n", "Ortho | \n", "rep1 | \n", "
Missing values are not measured zeros
\n", "Intensity plots generally omit missing measurements; completeness and\n", "detection plots count their absence. Measured zeros usually remain valid\n", "observations unless a function's zero-handling option says otherwise.\n", "The binary heatmap uses a detection threshold and needs an explicit\n", "missing-value encoding; the correlation heatmap requires complete input\n", "or a fill value. We explain those choices alongside their examples.\n", "Missing sample annotations are a separate issue: functions may reject,\n", "omit, or display them as a missing category. UpSet requires nonmissing\n", "category labels.
\n", "