Quaternary oxide composition generation#

This tutorial demonstrates how to generate a set of quaternary oxide compositions using the SMACT library and a modified smact_filter function from the smact_screening module then prepare the results for machine learning analysis.

Prerequisites#

Before starting, ensure you have the following libraries installed:

Open in Colab

# Install the required packages
try:
    import google.colab

    IN_COLAB = True
except:
    IN_COLAB = False

if IN_COLAB:
    !uv pip install smact[featurisers] --quiet

Workflow#

1. Import required libraries#

"""
This module imports necessary libraries and modules for generating and analyzing
quaternary oxide compositions using SMACT and machine learning techniques.

"""
# Standard library imports
from itertools import combinations, product

# Third-party imports
import pandas as pd
from matminer.featurizers import composition as cf
from matminer.featurizers.base import MultipleFeaturizer

# The parallel pool below comes from pathos, a SMACT dependency, rather than the standard
# library's multiprocessing. pathos serialises with dill, so it can send a function
# defined in a notebook cell, together with the globals that function refers to, out to
# worker processes. multiprocessing pickles such a function by reference instead, and
# fails with "AttributeError: Can't get attribute 'smact_filter' on <module '__main__'>"
# under the spawn start method, which is the default on macOS and Windows.
from pathos.multiprocessing import ProcessPool
from pymatgen.core import Composition

# Local imports
import smact
from smact import screening

"""
Imported modules:
- pathos: For parallel processing capabilities
- itertools: For generating combinations and products
- pandas: For data manipulation and analysis
- matminer: For materials data mining and feature extraction
- pymatgen: For materials analysis
- smact: For structure prediction and analysis of new materials
"""

2. Define elements and combinations#

Select the elements to use in our compositions

# Define the elements we are interested in.
#
# This tutorial runs on a small, hand-picked set of elements so the cells below
# finish in seconds rather than tens of minutes. The elements chosen span common
# solar-oxide-forming cation chemistry -- e.g. Cu2O, ZnO, TiO2, SnO2, Fe2O3 and
# Bi2O3 are all well-known photovoltaic or photocatalytic oxides -- so the demo
# still produces varied, chemically meaningful results.
#
# The published dataset behind this tutorial screens (almost) every element in
# the periodic table this way instead of just ten -- see "Reproducing results"
# at the end for what that involves.
all_el = smact.element_dictionary()
demo_symbols = ["Na", "Mg", "Al", "Ca", "Ti", "Fe", "Cu", "Zn", "Sn", "Bi"]
good_elements = [all_el[x] for x in demo_symbols]

# Generate all possible combinations of 3 elements from good_elements
all_el_combos = combinations(good_elements, 3)

3. Define SMACT filtering function#

Create a function to filter element combinations based on SMACT criteria:

def smact_filter(els):
    """
    Filter element combinations based on SMACT criteria.

    This function takes a combination of elements and applies SMACT
    (Semiconducting Materials from Analogy and Chemical Theory) tests
    to generate potential quaternary oxide compositions.

    Args:
        els (tuple): A tuple containing three Element objects.

    Returns:
        list: A list of tuples, each containing a set of elements and their ratios
              that pass the SMACT criteria.
    """
    all_compounds = []
    elements = [e.symbol for e in els] + ["O"]

    # Get Pauling electronegativities
    paul_a, paul_b, paul_c = (el.pauling_eneg for el in els)
    electronegativities = [paul_a, paul_b, paul_c, 3.44]  # 3.44 is for Oxygen

    # Iterate through all possible oxidation state combinations
    for ox_states in product(*(el.oxidation_states for el in els)):
        ox_states = list(ox_states) + [-2]  # Add oxygen's oxidation state

        # Test for charge balance
        cn_r = smact.neutral_ratios(ox_states, threshold=8)

        if cn_r:
            # Electronegativity test
            if screening.pauling_test(ox_states, electronegativities):
                compound = (elements, cn_r[0])
                all_compounds.append(compound)

    return all_compounds

4. Process element combinations#

Apply the SMACT filter to all element combinations in parallel:

For the small demo set above this finishes in a few seconds. The pool comes from pathos rather than the standard library’s multiprocessing, so that smact_filter as defined in the cell above can be sent to worker processes on every platform. See the note in the imports cell for details.

def process_element_combinations(all_el_combos):
    """
    Process all element combinations in parallel.

    This function applies the smact_filter to all element combinations
    using a pathos process pool to improve performance.

    Args:
        all_el_combos (iterable): An iterable of element combinations.

    Returns:
        list: A flattened list of all compounds that pass the SMACT criteria.
    """
    with ProcessPool() as p:
        # Apply smact_filter to all element combinations in parallel
        result = p.map(smact_filter, all_el_combos)
        # pathos keeps its pool in a module-level cache and leaving the with block does
        # not shut the workers down, so clear it to avoid stray processes outliving the
        # cell. A later ProcessPool() then starts a fresh pool rather than reusing this one.
        p.clear()

    # Flatten the list of results
    flat_list = [item for sublist in result for item in sublist]
    return flat_list


# Process all element combinations
flat_list = process_element_combinations(all_el_combos)

# Print the number of compositions found
print(f"Number of compositions: {len(flat_list)}")

5. Generate pretty formulas#

This step turns the generated compositions into pretty formulas, again in parallel. For the demo set above there should be a couple of thousand unique formulas.

def comp_maker(comp):
    """
    Convert a composition tuple to a pretty formula string.

    Args:
        comp (tuple): A tuple containing two lists - elements and their amounts.

    Returns:
        str: The reduced formula of the composition as a string.
    """
    # Create a list to store elements and their amounts
    form = []
    # Iterate through elements and their amounts
    for el, ammt in zip(comp[0], comp[1]):
        form.append(el)
        form.append(ammt)
    # Join all elements into a single string
    form = "".join(str(e) for e in form)
    # Convert to a Composition object and get the reduced formula
    pmg_form = Composition(form).reduced_formula
    return pmg_form


# Apply comp_maker to all compositions in flat_list, in parallel
with ProcessPool() as p:
    pretty_formulas = p.map(comp_maker, flat_list)
    # As above, clear the cached pool so its workers do not outlive this cell.
    p.clear()

# Create a list of unique formulas
unique_pretty_formulas = list(set(pretty_formulas))
# Print the number of unique composition formulas
print(f"Number of unique compositions formulas: {len(unique_pretty_formulas)}")

6. Create DataFrame and add descriptors#

Create a DataFrame from the unique formulas and add composition-based descriptors:

# Create a DataFrame from the unique pretty formulas
new_data = pd.DataFrame(unique_pretty_formulas).rename(columns={0: "pretty_formula"})

# Remove any duplicate formulas to ensure uniqueness
new_data = new_data.drop_duplicates(subset="pretty_formula")

# Display summary statistics of the DataFrame
# This will show count, unique values, top value, and its frequency
# new_data.describe()


# Add descriptor columns
# This will take a little time as we have over 1 million rows


def add_descriptors(data):
    """
    Add composition-based descriptors to the dataframe.

    This function converts formula strings to composition objects and calculates
    various features using matminer's composition featurizers.

    Args:
        data (pd.DataFrame): DataFrame containing 'pretty_formula' column.

    Returns:
        pd.DataFrame: DataFrame with added descriptor columns.
    """
    # Convert formula strings to composition objects
    data["composition_obj"] = data["pretty_formula"].map(Composition)

    # Initialize multiple featurizers
    feature_calculators = MultipleFeaturizer(
        [
            cf.Stoichiometry(),
            cf.ElementProperty.from_preset("magpie"),
            cf.ValenceOrbital(props=["avg"]),
            cf.IonProperty(fast=True),
            cf.BandCenter(),
            cf.AtomicOrbitals(),
        ]
    )

    # Set the number of parallel jobs. Adjust this as needed.
    feature_calculators.set_n_jobs(1)

    # Calculate features
    feature_calculators.featurize_dataframe(data, col_id="composition_obj")

    # If you need to use feature_labels later, uncomment the following line:
    # feature_labels = feature_calculators.feature_labels()

    return data


# Apply the function to add descriptors
new_data = add_descriptors(new_data)

7. Save results to a CSV file#

# Save as .csv file
new_data.to_csv("All_oxide_comps_dataframe_featurized.csv", chunksize=10000)

Reproducing results#

Sections 1-7 above are a fast, illustrative demo on ten hand-picked elements, not the exhaustive search behind the published dataset. That screens (almost) every element in the periodic table the same way – swap demo_symbols in Section 2 for every element smact.element_dictionary() returns (excluding ones without usable oxidation-state data, e.g. noble gases and most radioactive/synthetic elements) and rerun Sections 4-7 unchanged.

Expect that to take tens of minutes to a few hours depending on your hardware and produce on the order of ~1.1M unique compound formulas – deliberately left out of this notebook as a runnable cell, since a tutorial that takes 40 minutes to execute defeats the point of being a tutorial.