This repository provides training for setting up and using the following data stack, a collection of tools for working on data analytics projects. The intended audience is students in my courses at the Jon M. Huntsman School of Business at Utah State University, students that I’m mentoring on projects at the Analytics Solutions Center (ASC), and collaborators on research projects.
The data stack consists of:
- Positron as the code editor or integrated development environment
- Python for data wrangling, visualizations, and modeling
- Quarto for communicating results with presentations, reports, dashboards, etc.
- GitHub for version control, project management, and collaboration
Every modern data stack also includes AI tools. All Utah State students have access to a specific set. While AI can help learning and productivity (e.g., drafting and debugging code, explaining concepts in new ways, practicing for interviews), it can be harmful when we use it to replace rather than supplement thinking and decision-making—especially when we don’t know enough about a topic to evaluate what the AI generates. If you use AI tools, be thoughtful and transparent, including reviewing what the AI generates and citing the AI tool you use.
A code editor or integrated development environment (IDE) is the most important tool in the data stack. A good IDE provides a single tool for writing and running code, including communicating results and implementing version control. We’ll use Positron, a next-generation data science IDE. Built on VS Code’s open source core, Positron combines the multilingual extensibility of VS Code with essential data tools common to language-specific IDEs.
Get started by downloading and installing Positron and then walking through the following highlights of some of Positron’s essential features, especially its data-friendly functionality.
If you’ve used VS Code, Positron’s layout will look familiar. When selected from the vertical activity bar, the explorer on the left shows the folder you have open, which also establishes your working directory (i.e., the location on your machine for your project files). For students in my courses, create and open a folder that will serve as the home for all of your work, including course projects. The editor in the center is where you code. Two obvious differences are the console (in the bottom panel by default) and the session (on the right by default).
The console is where code runs (use Cmd/Ctrl + Enter to run the currently selected code in the editor) and is separate from the terminal (also called the command line or shell, in the bottom panel by default). The terminal is used to interact with your operating system to do things outside running code, like installing Python libraries. There is no console in VS Code and so code also runs in the terminal, which often means you have multiple terminals running for different purposes. The session displays the variables (e.g., data, functions, and methods) you’ve loaded and plots you’ve created.
You can click on the data frame icon to the right of any data you’ve loaded in the session to open the data explorer. The data explorer provides a summary of the data, including simple visualizations, and allows you to quickly sort and filter the data to inform data wrangling.
The data explorer is designed to facilitate coding, not replace it. If you want to implement any of the sorting, filtering, etc. you make using the data explorer in code, click the convert to code button in to the action bar at the top.
You can also click on Excel, comma-separated value (CSV), Parquet, PDF, and other files in your working directory to view them without needing to load them or use another program.
Along with variables, the session has a dedicated pane for visualizations, including a history gallery to click through and easily compare previous plots. Visualizations can also be opened as a separate tab in the editor pane. This includes support for interactive plots.
Including a question mark after most any function, method, or attribute in the console will open the help (on the right by default). The help serves as a built-in web browser to allow you to reference online documentation, including function and method details and examples you can copy and use.
Posit Assistant, selected from the vertical activity bar, is an AI tool integrated in Positron with contextual awareness of everything in your working directory. You can use Posit Assistant to ask questions, edit code, and function as an agent to accomplish specific tasks. It can use a variety of providers to interact with your data and code.
VS Code, and thus Positron, is highly extensible. Installed extensions are visible when selected from the vertical activity bar, along with the ability to search additional available extensions.
Since Positron is open source, there are certain proprietary VS Code extensions that aren’t available in Positron. The search functionality references the open source Open VSX Registry for all available extensions.
The command palette is the primary way to manage options (e.g., customize layout and themes) and is a mainstay of the shortcut-heavy VS Code. Open with Cmd/Ctrl + Shift + P.
In the upper-right corner are a set of icons to customize the layout of Positron, including a number of layout presets and toggles for side bars and panels.
There is more that Positron can do, including connecting to databases and remoting into virtual machines. Additionally, since Positron is built on VS Code’s open source core, VS Code’s excellent documentation remains largely relevant.
Python is a general purpose, open source programming language, often referred to as “the second-best language for everything.” Notably, Python comes pre-installed on some operating systems (OS). This version should not be used or modifed by anyone except the OS itself. For this and other reasons, you’ll need the ability to install and maintain multiple versions of Python on the same computer.
There are many ways to install and maintain Python versions. However, uv has emerged as the industry standard. Get started by opening the command palette in Positron (Cmd/Ctrl + Shift + P) and running “Python: Install Python via uv.”
This command installs uv and Python for you via the terminal. Note that the needed extension to run Python in Positron comes pre-installed. The following highlights of some of Python’s essential features for data analysis.
Most of the coding we do for data analysis uses functions, where the input is some data and the output is some transformation, visualization, or model results. However, Python is an object-oriented programming language, which creates a difference between functions and methods, a kind of function that only works for specific objects.
Python functions are typically namespaced, meaning the name of the
library they’re from is referenced with the function, as in
library.function(). Functions can also be namespaced with the
alias of the library that was set when the library was imported. For
example, using the Polars alias pl in
pl.read_csv('customer_data.csv') to read in customer_data.csv.
Methods are functions nested within object types and are namespaced
with an object name of the given type as in object.method(). For
example, using the Polars select method in
customer_data.select(pl.col('income')) to select the income column
in the customer_data DataFrame object.
Besides functions and methods, attributes are object-specific
features and are, like methods, namespaced with an object name of the
given type as in object.attribute, but without any parentheses. For
example, the column names of the customer_data DataFrame object can be
referenced with customer_data.columns.
Python is a big tent and there are a lot of different libraries or packages that have been developed by users to facilitate coding for different applications. Each library or package is a set of functions, methods, documentation, and sometimes data. While there are many Python libraries, I recommend the following for the three primary data analytics tasks.
- Data Wrangling: Polars is a fast, self-consistent library for data wrangling (i.e., cleaning and manipulating data) that is growing in popularity as an alternative to pandas. Additionally, when you read in data to wrangle, be sure to write relative file paths using pyhere.
- Visualizations: Plotnine is a library built using the consistency of the grammar of graphics philosophy for visualizations. It is a port of R’s {ggplot2} package, which is the industry standard across open source programming languages.
- Modeling: scikit-learn is the most widely used library for machine learning, but it doesn’t do statistical inference. For statistical inference, the statsmodels and Bambi libraries are used for frequentist and Bayesian modeling, respectively.
Just like you only need to install Positron, uv, and Python once, you only need to install Python libraries once. You can see what libraries or packages you have installed by selecting packages from the vertical activity bar.
This list includes libraries you’ve installed, the libraries that come pre-installed with Python, and any dependencies or the libraries that those libraries depend on. You can also see which libraries have newer versions you can update to.
In order for our code to be reproducible, we need to maintain project environments. A project environment is composed of both Python and the libraries (including the dependencies) used for a given project. What makes a project environment reproducible is keeping track of the Python and library versions we’re using for a given project so that it can be easily reproduced on another machine by you (including future you) or someone else. It’s easy to set up and manage a project environment with uv.
- Open the project working directory (e.g., the course folder) in Positron. Note that this working directory should not be in a location on your local machine that is being synced to the cloud via OneDrive, iCloud, etc.
- Run
uv initin the terminal to initialize a project environment. This creates apyproject.tomlfile with metadata about the project and a hidden.python-versionfile that specifies the default version of Python for the project. (It also createsmain.pyandREADME.mdfiles that you can use or delete.) - With the project environment initialized, you can install libraries.
For example, running
uv add polarsvia the terminal installs Polars (and any dependencies) and creates both auv.lockfile that keeps track of the versions of the libraries you’ve installed and a hidden/.venvreproducible (or virtual, hence the “v” in venv) environment folder that serves as the project library.
Note
The terminal (i.e., the command line or shell) is the programming interface into your OS itself. Note that the name of the terminal will be different based on your OS. The macOS terminal is Zsh, the Linux terminal is Bash, and the Windows terminal is PowerShell.
You can think of the console as a specialized terminal for running code only while the general terminal is where we can interact with the operating system for everything outside of running code. This includes uv and Python, which Positron handled for us with the “Install Python via uv” command, and now setting up a project environment and installing Python libraries.
Note that to use something you just installed via the terminal, you often need to first restart or create a new terminal.
All Python libraries are installed in a single, global library on your
computer known as the system library. The fact that we have a
project library highlights an important feature of making project
environments reproducible: Each project will have its own project
library and thus be isolated. If two projects use different versions of
the same package, they won’t conflict with each other because they’ll
each have their own project library. (Well, not exactly. Python employs
a global cache to avoid having to install the same version of a given
library more than once. The project library will reference the global
cache.) Whenever you install new libraries for your project, the
uv.lock file is automatically updated.
There is a lot more that uv can do. You
can manually install specific versions of Python, such as
uv python install 3.13.4 to install Python 3.13.4, and view Python
versions that are available to install with uv python list. If someone
is using another tool to install libraries instead of uv (e.g., pip),
they will likely need a requirements.txt file or a pylock.toml file
in place of the pyproject.toml and uv.lock files to reproduce the
project environment, which you can generate for them with
uv export --format requirements.txt or uv export -o pylock.toml,
respectively.
Just like different application use different Python libraries, so do different applications lend themselves to different coding styles. For example, the code that runs the transaction backend for the payment system of an online application needs to be incredibly efficient and secure since its managing sensitive information and being used at a high frequency. On the other hand, the code for a data analysis with moderate-sized data, even if the resulting report is reproduced often, runs very infrequently so efficiency takes a back seat to the code being readable.
Code used in production must be efficient, but often at the cost of it being readable—especially when custom functions are created. Code used in data analysis tends to be custom and focused more on readability, in part because it often doesn’t need to be efficient. This is also important if you are coding with the help of an AI tool, which will naturally gravitate toward needless efficiency and an overabundance of code I refer to as AI bloat. As you code for data analysis, focus on readability. You’ll be more productive when working with AI tools since you should be able to better understand the output.
Perhaps the most readable code is produced with a technique called method chaining. Instead of saving out intermediate objects for every step in a set of method calls, just chain them all together.
The resulting code, which has to be enclosed in parentheses, can be read like a sentence as we move from one method to the next. Using Cmd/Ctrl + Enter to run the currently selected code in the editor will run an entire method chain. Method chaining is enabled by Polars syntax and is mirrored in the composition of Plotnine’s grammar of graphics.
Finally, the style of the code (e.g., spacing and indenting conventions) varies wildly from person to person, which can cause needless friction when collaborating, much like not using relative file paths. Positron comes pre-installed with Ruff, a code linter that will enforce standard style and formatting.
Tip
Python might be the most commonly used open source programming language for data wrangling, visualizations, and modeling—but it’s not the only one. The three most popular languages for data analytics are Julia, Python, and R (the Jupyter kernel was named for and designed to support all three). Each language comes with its own tradeoffs, culture, and overall vibe.
- Julia is the newest and fastest and was developed by mathematicians.
- Python is the most popular and diverse in terms of libraries and applications and was developed by computer scientists.
- R is the most narrowly focused on data analytics and culturally cohesive and was developed by statisticians.
If you want to learn more than one of the three languages, and you arguably should, I recommend focusing on becoming proficient in one language first and then transferring that understanding to picking up a second. For example, see how I learned Python coming from a background using R.
Much of the code we write for a project can use flat text Python .py
scripts. However, if we need to produce an output in a format other than
code, for example a report, then we should use
Quarto, an open source publishing system where
we can combine writing along with code and its output. If you’ve used
Jupyter notebooks, Quarto documents will be familiar with designated
sections for writing and code.
Like in Jupyter notebooks, we can run each of the code blocks individually or all at once with the play buttons. (To have the output appear in the Quarto document itself like in Jupyter notebooks rather than the console only, toggle on Preferences > Settings > Quarto > Inline Output: Enabled.) Unlike Jupyter notebooks, the code blocks in Quarto documents are flat text Python scripts so we can still use Cmd/Ctrl + Enter to run individual lines of code within a code block.
The most important difference is that the notebook format of a Quarto document is simply a means to an end. Quarto can take whatever we produce within the document and render it into a Word document, PowerPoint presentation, PDF, Revealjs slide deck, interactive dashboard, website, etc. Browse through the gallery to see what sort of things are possible.
The Quarto extension comes pre-installed with Positron. Whenever you make a change to a Quarto document, render the document (click on Preview or use Cmd/Ctrl + Shift + K) into its specified format and a preview of the rendered document will appear in Positron’s viewer (in the right pane by default).
If you are using Python within the Quarto document, Quarto will render
the output using the Jupyter kernel in the background. In fact, as
needed, a Quarto document can be used in conjunction with a Jupyter
notebook to
render into all of these different outputs. For example, we can render
pydata.ipynb into a PDF using quarto render pydata.ipynb --to typst
in the terminal.
Quarto’s documentation is comprehensive. The following sections highlight some of the essential features of Quarto documents.
The header of any Quarto document is coded in YAML (i.e., Yet Another
Markup Language), which follows a simple key: value syntax. To render
documents into a PDF, use format: typst.
Typst is modern,
fast typesetting software for creating PDFs and comes pre-installed with
Quarto. For GitHub documents, use format: gfm. When you render your
Quarto document, it will create a separate markdown document using
GitHub Flavored Markdown that GitHub can parse as HTML.
Quarto documents use markdown, just like in Jupyter notebooks. Markdown is a simple, generic typesetting syntax. Note that GitHub also recognizes this syntax, including in issues and pull requests.
Sometimes working with markdown alone can be challenging. Positron includes a visual mode you can access inside any Quarto document. The visual mode includes some point-and-click options to help you produce markdown syntax, which can be especially helpful for things like tables and citations.
Quarto allows us to include code blocks and output as part of the document. Much like Jupyter notebooks, you can include Julia, Python, and R code as well as C++, Stan, and other code blocks and output.
There are a variety of options for each code block. In addition to
specifying the language used within the code block, the code block can
be given an identifier, can have warnings suppressed, can run without
producing output, etc. These options are specified using YAML syntax
following the hashpipe operator #| within the body of the code block.
Any code block YAML that should apply to the document in its entirety
can simply be moved into the header YAML.
If you need to include any math, you shouldn’t be surprised that there’s
a typesetting syntax for that. It’s tied to
LaTeX (pronounced “lah-tech” or
“lay-tech”) and our primary interest is using it’s math
syntax. Use
$ around any in-line LaTeX notation or $$ around equations specified
as a separate line. For example, we can reference
Git is a powerful system for version control (also called source control). It is the industry standard for software development and has been adopted to provide structure for data analysis as well. GitHub is an online hosting service where each project lives in its own repository or repo. Learning to use Git and GitHub not only aids in collaboration, it will ultimately allow you to develop an online portfolio of work.
Get started by signing up for a GitHub account using a professional username and your student email address and downloading and installing Git. You’ll also need to introduce yourself to Git by modifying and running the following in the terminal, where the email address is the same you used for your GitHub account.
git config --global user.name "Your Name"
git config --global user.email "your.email@example.com"
For ASC and research projects, I use my project template and I recommend using it for course project work as well. Click on “Use this template” and create a new repository with a short, lowercase, hyphenated slug as the repository name that’s consistent with the project (e.g., advanced-coursework).
One person will maintain the repository and have it connected to their account (e.g., the mentor for ASC projects) while others working on the project can be added as collaborators. Anyone with access to the project repository can copy (i.e., fork) it to save, maintain, or contribute to if they aren’t a collaborator.
Note that there are certain limitations to the size and type of files that can be hosted (i.e., pushed to GitHub). There are also certain things that shouldn’t be accessible by the public (e.g., private data). For these reasons, we have files and folders that are pushed to GitHub and those that are not. Here’s how the project repository is organized:
/codeScripts with prefixes (e.g.,01_import-data.py,02_clean-data.py) and functions in/code/src./dataSimulated and real data, the latter not pushed./figuresPNG images and plots./outputOutput from model runs, not pushed./presentationsPresentation slides./privateA catch-all folder for miscellaneous files, not pushed./writingPaper, report, and case studies./.quartoHidden Quarto project library, not pushed./.venvHidden Python project library, not pushed..gitignoreHidden Git instructions file..python-versionHidden Python version file.LICENSEMIT License for “as is” permission.README.mdGitHub-flavored markdown README rendered fromREADME.qmd.README.qmdQuarto markdown README to edit and render intoREADME.md._quarto.ymlQuarto project configuration file.pyproject.tomlPython project environment configuration file.uv.lockPython project environment lockfile.
Any file or folder that begins with a period is hidden (i.e., you won’t
see it in your OS file explorer by default, but you will see it in the
explorer in Positron). The .gitignore file is what controls which
files and folders are pushed. Note that the project
template also includes
instructions for using the project environment and has a link to this
training repo.
To stay organized, manage your project by keeping track of tasks using GitHub’s issues (see the tab at the top of the project repository on GitHub). There you can have an ongoing conversation and close tasks out when a given issue is completed or resolved.
Be sure to tag collaborators you want to see a specific comment (e.g.,
@marcdotson). Think of this as an email thread or chat channel except
all of the conversations are in one place, easily searchable, and
automatically archived as part of the version control.
Once you have created or have access to a project repository, you can
clone it. Cloning simply means you’re creating a local copy of the
repository, though the clone isn’t just a copy of the folder. Git still
works in the background keeping track of changes and managing the
version control. After connecting Positron and GitHub, open the command
palette in Positron, use the Git: Clone command, and select the
project repository you’d like to clone.
As detailed in the project template README, instead of having to create
a project environment, you simply need to open the project working
directory in Positron and run uv run in the terminal. This will
install the correct .python-version if you don’t have it, create the
hidden /.venv project library, and install the correct versions of the
needed libraries as specified in the uv.lock file.
You only need to clone the project repository one time. Please note that
any of the files and folders that aren’t pushed to GitHub can be created
in your cloned repository without any impact on the repository hosted on
GitHub. For example, I typically store PDFs and other related materials
that I don’t want to (or can’t) share in the /private folder so that
the resources I need are all within the same directory on my machine.
Remember that Git and GitHub are built for software development.
Following that analogy, Git operates through the use of
branches.
Each branch in a repository is a separate version of the repository that
exists in parallel and is focused on a specific issue. For example, we
could have a branch called initial-model and another branch called
data-cleaning. For version control to be useful, we need to be concise
and descriptive with branch and other naming conventions (i.e., no
final-final-draft-02 nonsense here).
Every repository has a main branch. If this were software, the main
branch would be the branch that is being used in production. Never make
changes directly to the main branch. You can see which branch you’re
working in by looking at the bottom left corner in Positron. Assuming
the branch you need to work on has already been created, the first thing
you should do when starting to work is navigate to the branch you want
to work in. Use Git: Checkout to... via the command palette to select
the correct branch.
You’ve identified what you need to work on using issues, cloned the project repository, and made sure you’re working on the correct branch. Finally, you can get to work! Then what? Once you’ve made a number of changes to your cloned project respository, how do you share that with your collaborators?
If you click on source control from the vertical activity bar, you’ll see all of the files you’ve changed.
You first need to stage these changes using the plus sign next to the files. It’s best practice to stage changes across files that relate to the specific work you’ve done and want to track within the version control graph.
Once the files are staged, you need to write a message that describes what you’ve done with the staged files. Like the branch names, these should be short and descriptive, like “Cleaned up errors to the final model.” It is the branch names and these commit messages that provide a record of the work we have done.
With a descriptive message, you are ready to commit. This is like saving a file, except we can save multiple files all at once associated with the commit message we’ve written.
These changes are now archived as part of the version control on our cloned project repository. To share them with our collaborators, we need to sync them with the repository on GitHub.
Technically, sync executes a push of our changes to the repository on GitHub while also executing a pull of any changes our collaborators have previously pushed to GitHub. Alongside the branch name in the bottom left you should see how many commits you have to pull and push at any given time.
To summarize, your daily work in Positron will look like this:
- Open the cloned project repository in the explorer pane to set your working directory.
- Make sure you’ve checked out the correct branch to work in (the branch name you’ve checked out is in the bottom left corner).
- Pull changes that have been pushed since you last worked on the project using the sync button in source control or next to the branch name in the bottom left corner.
- Once you’ve done some amount of work, stage and commit changes with a descriptive message in the source control pane.
- Push changes using the sync button in the source control pane or next to the branch name in the bottom left corner.
When you’ve completed work on the issue associated with the branch, you can create a pull request on GitHub.
A pull request is exactly what it sounds like—a request submitted to the
repository maintainer to pull the changes you’ve made on your branch
into main. This allows the maintainer, or someone they assign, to
review what you’ve done, have a conversation with you about it as part
of the pull request itself (which looks a lot like an issue tied
specifically to the pull request), and eventually pull what you’ve done
into main. After the pull request is completed, the branch specific to
that issue can be deleted and the associated issue can be closed out.
Note that when a branch is deleted on GitHub, it will still exist in
your cloned repository. This isn’t necessarily a problem, though if you
commit changes to a closed branch it will force the branch open again.
Remember to make sure you’re working on the correct branch. Eventually
you may want to clean up branches that have been merged into main and
closed on GitHub by using Git: Delete Branch... via the command
palette, followed by the branch name. You may also need to use
Git: Fetch to prune tracking branches that are no longer on remote
(i.e., on GitHub).























