I am getting started into Astronomy datasets. One of the first things I wanted to do is get adjusted and acquainted with TensorFlow. Since I have an Apple M1, I wanted to get tensorflow installed and using the integrated libraries so when the datasets are written, they are more native using the Apple Metal framework.

Here are the scratch notes.

  1. Install Xcode Command Line Tools
  2. Install Miniforge
  3. Install Tensorflow and its dependencies
  4. Install Jupyter Notebook, Pandas
  5. Run a Benchmark by training the MNIST dataset

Step 1: Xcode Command Line Tools. I have already installed these on my mac. If it is not already installed on your system, run:

xcode-select --install

Step 2. Install Miniforge. Install miniforge for arm64 (Apple Silicon) from miniforge GitHub.

Miniforge matters here and it is worth saying why rather than just doing it. Anaconda’s own installer was x86_64 for a long time, so a conda environment created with it ran every package through Rosetta 2 and none of the native Metal path was reachable. Miniforge is built for arm64 and defaults to conda-forge, which had native Apple Silicon builds of the numeric stack long before the main channel did. Getting this wrong is the single most common reason people end up with a working install that is inexplicably slow: they are running an emulated x86 Python on an ARM machine.

By default miniforge activates a base environment in every shell. To turn that off:

conda config --set auto_activate_base false

Step 3. Install TensorFlow. Install the dependencies first:

conda install -c apple tensorflow-deps

That package pulls the numeric libraries TensorFlow links against, built for arm64 and pinned to versions the Apple build expects. Then base TensorFlow and the GPU plugin:

pip install tensorflow-macos

pip install tensorflow-metal

Those two are a matched pair. tensorflow-metal registers itself through TensorFlow’s PluggableDevice interface and routes supported kernels onto Metal Performance Shaders; it is built against a specific tensorflow-macos version, and a mismatch shows up as an import-time symbol error rather than a useful message. Install them together and pin them together.

Step 4. Install Jupyter Notebook & Pandas

conda install -c conda-forge -y pandas jupyter

Step 5. Run a Benchmark by training the MNIST dataset. Install TensorFlow Datasets:

pip install tensorflow_datasets

Make sure the conda environment is activated, run jupyter notebook, and create a new Python 3 notebook. First confirm the device is visible:

import tensorflow as tf
print("Num GPUs Available: ", len(tf.config.list_physical_devices('GPU')))
print(tf.__version__)

A count of 1 means the plugin loaded. It does not mean your model is using it, which is a different question and a more important one. To see where the work actually lands:

tf.debugging.set_log_device_placement(True)

That prints the chosen device per operation. Anything the Metal plugin does not implement falls back to the CPU silently, so this is how you find the layer that is quietly costing you the speedup.

The benchmark code snippet I used came from TensorFlow Issues. Copy it into the notebook and examine the results. Worth checking the architecture while you are in there, since it catches the Rosetta case immediately:

import platform; print(platform.machine())   # want: arm64