Quaternary oxide composition generation#
This tutorial demonstrates how to generate a set of quaternary oxide compositions using the SMACT library and a modified smact_filter function from the smact_screening module then prepare the results for machine learning analysis.
Prerequisites#
Before starting, ensure you have the following libraries installed:
# Install the required packages
try:
import google.colab
IN_COLAB = True
except:
IN_COLAB = False
if IN_COLAB:
!uv pip install smact[featurisers] --quiet
Workflow#
1. Import required libraries#
"""
This module imports necessary libraries and modules for generating and analyzing
quaternary oxide compositions using SMACT and machine learning techniques.
"""
# Standard library imports
from itertools import combinations, product
# Third-party imports
import pandas as pd
from matminer.featurizers import composition as cf
from matminer.featurizers.base import MultipleFeaturizer
# The parallel pool below comes from pathos, a SMACT dependency, rather than the standard
# library's multiprocessing. pathos serialises with dill, so it can send a function
# defined in a notebook cell, together with the globals that function refers to, out to
# worker processes. multiprocessing pickles such a function by reference instead, and
# fails with "AttributeError: Can't get attribute 'smact_filter' on <module '__main__'>"
# under the spawn start method, which is the default on macOS and Windows.
from pathos.multiprocessing import ProcessPool
from pymatgen.core import Composition
# Local imports
import smact
from smact import screening
"""
Imported modules:
- pathos: For parallel processing capabilities
- itertools: For generating combinations and products
- pandas: For data manipulation and analysis
- matminer: For materials data mining and feature extraction
- pymatgen: For materials analysis
- smact: For structure prediction and analysis of new materials
"""
2. Define elements and combinations#
Select the elements to use in our compositions
# Define the elements we are interested in.
#
# This tutorial runs on a small, hand-picked set of elements so the cells below
# finish in seconds rather than tens of minutes. The elements chosen span common
# solar-oxide-forming cation chemistry -- e.g. Cu2O, ZnO, TiO2, SnO2, Fe2O3 and
# Bi2O3 are all well-known photovoltaic or photocatalytic oxides -- so the demo
# still produces varied, chemically meaningful results.
#
# The published dataset behind this tutorial screens (almost) every element in
# the periodic table this way instead of just ten -- see "Reproducing results"
# at the end for what that involves.
all_el = smact.element_dictionary()
demo_symbols = ["Na", "Mg", "Al", "Ca", "Ti", "Fe", "Cu", "Zn", "Sn", "Bi"]
good_elements = [all_el[x] for x in demo_symbols]
# Generate all possible combinations of 3 elements from good_elements
all_el_combos = combinations(good_elements, 3)
3. Define SMACT filtering function#
Create a function to filter element combinations based on SMACT criteria:
def smact_filter(els):
"""
Filter element combinations based on SMACT criteria.
This function takes a combination of elements and applies SMACT
(Semiconducting Materials from Analogy and Chemical Theory) tests
to generate potential quaternary oxide compositions.
Args:
els (tuple): A tuple containing three Element objects.
Returns:
list: A list of tuples, each containing a set of elements and their ratios
that pass the SMACT criteria.
"""
all_compounds = []
elements = [e.symbol for e in els] + ["O"]
# Get Pauling electronegativities
paul_a, paul_b, paul_c = (el.pauling_eneg for el in els)
electronegativities = [paul_a, paul_b, paul_c, 3.44] # 3.44 is for Oxygen
# Iterate through all possible oxidation state combinations
for ox_states in product(*(el.oxidation_states for el in els)):
ox_states = list(ox_states) + [-2] # Add oxygen's oxidation state
# Test for charge balance
cn_r = smact.neutral_ratios(ox_states, threshold=8)
if cn_r:
# Electronegativity test
if screening.pauling_test(ox_states, electronegativities):
compound = (elements, cn_r[0])
all_compounds.append(compound)
return all_compounds
4. Process element combinations#
Apply the SMACT filter to all element combinations in parallel:
For the small demo set above this finishes in a few seconds. The pool comes from
pathos rather than the standard library’s multiprocessing, so that smact_filter
as defined in the cell above can be sent to worker processes on every platform. See
the note in the imports cell for details.
def process_element_combinations(all_el_combos):
"""
Process all element combinations in parallel.
This function applies the smact_filter to all element combinations
using a pathos process pool to improve performance.
Args:
all_el_combos (iterable): An iterable of element combinations.
Returns:
list: A flattened list of all compounds that pass the SMACT criteria.
"""
with ProcessPool() as p:
# Apply smact_filter to all element combinations in parallel
result = p.map(smact_filter, all_el_combos)
# pathos keeps its pool in a module-level cache and leaving the with block does
# not shut the workers down, so clear it to avoid stray processes outliving the
# cell. A later ProcessPool() then starts a fresh pool rather than reusing this one.
p.clear()
# Flatten the list of results
flat_list = [item for sublist in result for item in sublist]
return flat_list
# Process all element combinations
flat_list = process_element_combinations(all_el_combos)
# Print the number of compositions found
print(f"Number of compositions: {len(flat_list)}")
5. Generate pretty formulas#
This step turns the generated compositions into pretty formulas, again in parallel. For the demo set above there should be a couple of thousand unique formulas.
def comp_maker(comp):
"""
Convert a composition tuple to a pretty formula string.
Args:
comp (tuple): A tuple containing two lists - elements and their amounts.
Returns:
str: The reduced formula of the composition as a string.
"""
# Create a list to store elements and their amounts
form = []
# Iterate through elements and their amounts
for el, ammt in zip(comp[0], comp[1]):
form.append(el)
form.append(ammt)
# Join all elements into a single string
form = "".join(str(e) for e in form)
# Convert to a Composition object and get the reduced formula
pmg_form = Composition(form).reduced_formula
return pmg_form
# Apply comp_maker to all compositions in flat_list, in parallel
with ProcessPool() as p:
pretty_formulas = p.map(comp_maker, flat_list)
# As above, clear the cached pool so its workers do not outlive this cell.
p.clear()
# Create a list of unique formulas
unique_pretty_formulas = list(set(pretty_formulas))
# Print the number of unique composition formulas
print(f"Number of unique compositions formulas: {len(unique_pretty_formulas)}")
6. Create DataFrame and add descriptors#
Create a DataFrame from the unique formulas and add composition-based descriptors:
# Create a DataFrame from the unique pretty formulas
new_data = pd.DataFrame(unique_pretty_formulas).rename(columns={0: "pretty_formula"})
# Remove any duplicate formulas to ensure uniqueness
new_data = new_data.drop_duplicates(subset="pretty_formula")
# Display summary statistics of the DataFrame
# This will show count, unique values, top value, and its frequency
# new_data.describe()
# Add descriptor columns
# This will take a little time as we have over 1 million rows
def add_descriptors(data):
"""
Add composition-based descriptors to the dataframe.
This function converts formula strings to composition objects and calculates
various features using matminer's composition featurizers.
Args:
data (pd.DataFrame): DataFrame containing 'pretty_formula' column.
Returns:
pd.DataFrame: DataFrame with added descriptor columns.
"""
# Convert formula strings to composition objects
data["composition_obj"] = data["pretty_formula"].map(Composition)
# Initialize multiple featurizers
feature_calculators = MultipleFeaturizer(
[
cf.Stoichiometry(),
cf.ElementProperty.from_preset("magpie"),
cf.ValenceOrbital(props=["avg"]),
cf.IonProperty(fast=True),
cf.BandCenter(),
cf.AtomicOrbitals(),
]
)
# Set the number of parallel jobs. Adjust this as needed.
feature_calculators.set_n_jobs(1)
# Calculate features
feature_calculators.featurize_dataframe(data, col_id="composition_obj")
# If you need to use feature_labels later, uncomment the following line:
# feature_labels = feature_calculators.feature_labels()
return data
# Apply the function to add descriptors
new_data = add_descriptors(new_data)
7. Save results to a CSV file#
# Save as .csv file
new_data.to_csv("All_oxide_comps_dataframe_featurized.csv", chunksize=10000)
Reproducing results#
Sections 1-7 above are a fast, illustrative demo on ten hand-picked elements, not
the exhaustive search behind the published dataset. That screens (almost) every
element in the periodic table the same way – swap demo_symbols in Section 2 for
every element smact.element_dictionary() returns (excluding ones without usable
oxidation-state data, e.g. noble gases and most radioactive/synthetic elements) and
rerun Sections 4-7 unchanged.
Expect that to take tens of minutes to a few hours depending on your hardware and produce on the order of ~1.1M unique compound formulas – deliberately left out of this notebook as a runnable cell, since a tutorial that takes 40 minutes to execute defeats the point of being a tutorial.