API Reference

MOF Structure

High level entry point to mofstructure.

MOFstructure wraps a single crystal structure and exposes the analyses that are normally wanted together: guest removal, porosity, deconstruction into building units, topology and open metal sites. Everything is computed on the guest free system, so a structure only has to be read once.

from mofstructure import structure

mof = structure.MOFstructure(filename=’UiO-66.cif’) pores = mof.get_porosity() metal_sbus, organic_sbus = mof.get_sbu()

The individual algorithms live in mofdeconstructor, porosity, generate_cgd and graph_net, and can be called directly when more control is needed. Topology is identified in Python; no Java runtime is involved.

class mofstructure.structure.MOFstructure(ase_atoms=None, filename=None)[source]

Bases: object

A single framework and the analyses that can be run on it.

Give it either an ASE atoms object or a path to any file ASE can read. The methods return their results rather than writing them, so they can be mixed freely:

mof = MOFstructure(filename=’UiO-66.cif’) pores = mof.get_porosity() metal_sbus, organic_sbus = mof.get_sbu() topology = mof.get_topology()

Guest molecules are stripped internally before each analysis, so a structure that still contains solvent does not need cleaning up first.

parameters:
  • ase_atoms: ASE atoms object, used in preference to filename

  • filename: path to a cif or any other ASE readable structure file

remove_guest()[source]

Simple function to remove guest molecules in porous system. Note that this can work for any periodic system.

returns:

ase_atoms (ase.Atoms): atom object with guest removed.

get_sbu(wrap_system=True, cheminfo=True, add_dummy=False)[source]

Extract the metal and linker secondary building units from the guest free system.

parameter

wrap_system (bool): If True, removes the effects of periodicity by merging the system. cheminfo (bool): If True, computes cheminformatic identifiers such as SMILES, InChI, and InChIKey. add_dummy (bool): If True, adds dummy atoms at the points of extension.

returns:

metal_sbu (list): A list of unique metal secondary building units linker_sbu (list): A list of unique organic secondary building units

get_ligands(wrap_system=True, cheminfo=True, add_dummy=False)[source]

Extract the metal and linker secondary building units from the guest free system.

parameter:

wrap_system (bool): If True, removes the effects of periodicity by merging the system. cheminfo (bool): If True, computes cheminformatic identifiers such as SMILES, InChI, and InChIKey. add_dummy (bool): If True, adds dummy atoms at the points of extension.

returns:

metal_clusters (list): A list of unique metal atoms or clusters. organic_ligands (list): A list of unique organic ligands.

get_porosity(probe_radius=1.86, number_of_steps=10000, rad_file=None, high_accuracy=True, timeout=1800)[source]

A function to compute porosity data for a system.

parameters:

  • probe_radius: float

    Radius of the probe (default: 1.86).

  • number_of_steps: int

    Number of GCMC simulation cycles (default: 10000).

  • high_accuracy: bool

    If True, perform high-accuracy computations.

  • rad_file: str

    Optional file of user defined atom radii. Must have the .rad extension.

  • timeout: float

    Seconds allowed for one structure before zeo++ is killed. A structure that runs past it returns the record below with None in place of each number.

returns:
pore: dict

The same keys whatever happens, so a directory of structures gives rows of one shape:

  • av_volume_fraction: Accessible volume void fraction.

  • av_a3: Accessible volume in A^3.

  • asa_a2: Accessible surface area in A^2.

  • asa_m2_per_cm3: Accessible surface area in m2/cm3.

  • number_of_channels: Number of channels, that is pores, present in the system.

  • lcd_a: The largest cavity diameter, the diameter of the largest sphere that can be inserted into the porous system without overlapping any atoms.

  • lfpd_a: The largest included sphere along the free sphere path, that is the largest sphere that can be inserted into the pore.

  • pld_a: The pore limiting diameter, the largest sphere that can freely diffuse through the porous network without overlapping any atoms.

  • porosity_status: ‘ok’, ‘timeout’ or ‘failed:<code>’.

get_topology(method='all_node', *, decimals=<object object>, include_edge_centers=<object object>, fallback_to_input_cgd=False, timeout=300, refine_cgd=False)[source]

Compute topology information for the guest-free system.

Identification is done in Python by mofstructure.topology. It was validated against Systre on 196 real frameworks across three corpora without a single disagreement, and against a third implementation, CrystalNets, on a further 111.

Two fields have changed meaning and the old ones are kept beside them rather than being quietly redefined:

topology_hash was a digest of Systre’s relaxed geometry, rounded coordinates and cell. It is now the digest of the canonical key, which is a stronger identifier: the same net written in a supercell, with its atoms reordered or its origin moved, hashes the same, where the geometric digest does not. Values stored under the old scheme will not match, which is why key_version travels with it.

cgd is now written from the ideal embedding this package derives from the topology rather than from Systre’s relaxation. Both describe the same net; they differ in setting, since Systre reports a conventional cell and this reports the primitive one, so a cubic a=4.899 there appears here as a=4.492 with 45 and 60 degree angles.

parameters:
  • method: str

    Deconstruction to use; any that TopologyExtractor.build_cgd accepts, or “auto” to choose from the structure.

  • decimals: int

    Deprecated and ignored, removal planned. The key is exact and integral, so the digest no longer depends on a rounding choice.

  • include_edge_centers: bool

    Deprecated and ignored, removal planned; the embedding writes no edge centres.

  • fallback_to_input_cgd: bool

    If True, return the deconstruction’s own CGD when no ideal embedding can be built, which happens for a net that is not 3-periodic.

  • timeout: int or None

    Seconds allowed for identification, None for no limit.

  • refine_cgd: bool

    Write the CGD from an embedding refined for uniform edge lengths rather than from the exact barycentric one. The exact embedding is reproducible and belongs beside the key; the refined one is what a builder such as AuToGraFS wants, since a spread of two between longest and shortest edge means a linker cannot span both.

returns:
python dictionary
Mapping containing:
  • topology, or None when no archive names the net. mofstructure.topology.analyse reports that case as UNNAMED_TOPOLOGY instead, so that a stored column is never empty; here a caller tests for None.

  • dimension

  • td10, the coordination sequence over ten shells summed with the vertex itself, averaged over vertex orbits and rounded. None when no ideal embedding could be built.

  • topology_hash, the same value as key_hash under the name earlier releases used

  • cgd, None on the same terms as td10

  • key, key_hash, key_version

  • topology_source, names

  • status, and detail when status is not “ok”

get_ligand_cluster_fingerprint()[source]

Describe how the ligands meet the metal clusters.

Where get_topology names the net, this reads the deconstruction directly, so it answers for every framework, including the ones no archive names and the ones that have no stable net at all. It counts each ligand and cluster species, how many clusters each ligand bridges, and with what denticity, which is what makes it sensitive to defects: a missing linker lowers a cluster’s connectivity, a linker hanging by one end is listed under terminal alongside its formula, and a carboxylate that has dropped from bridging to monodentate shows in the denticity histogram. It does not change when the atoms are listed in a different order, when the cell origin moves, or when the same crystal is given as a supercell.

returns:
python dictionary
Mapping containing:
  • clusters

  • ligands

  • terminal

  • refinement

  • formula_units

  • fingerprint_hash

draw_topology(method='all_node', supercell=(1, 1, 1), filename=None, show=False, show_structure=True, show_unit_cell=True, show_linker_sbu=True, show_topology=False)[source]

Draw the complete topological network over the real framework geometry.

Each node sits at the real centroid of the atoms it represents (metal atoms/clusters and the carboxyl and linker vertices), and edges follow the framework connectivity, including bonds that cross the cell. The figure is presented as a molecular structure rather than an axis-based plot: rotate, zoom and hover a node for its coordination. Nodes in neighbouring periodic images are included whenever an edge reaches them, so every visible edge has a visible node at both ends.

parameters:
  • method: str

    Node definition: “sbus”, “all_node”, “single_node” or “ligand_cluster”.

  • supercell: tuple of three ints

    How many cells to draw along a, b, c. (1, 1, 1) draws one cell plus the edges leaving it.

  • filename: str, optional

    If given, write the figure to this path. An .html file stays interactive; other extensions (.png, .pdf, …) need the optional ‘kaleido’ package.

  • show: bool

    If True, open the figure in a browser.

  • show_structure: bool

    If True, draw the framework atoms and bonds behind the net.

  • show_unit_cell: bool

    If True, draw the boundary of the displayed unit cells.

  • show_linker_sbu: bool

    If True, show the centre-to-centre network produced by the selected topology method. Its nodes and connections therefore change when the method changes.

  • show_topology: bool

    If True, also show the abstract topology nodes and blue edges. It is False by default so the SBU-linker connectivity remains visually unambiguous.

returns:

plotly.graph_objects.Figure

get_oms()[source]

Function to compute open metal sites

mofstructure.structure.get_coordination_environment(coordination_spheres, metal_element, coordination_number)[source]

Returns the coordination environment (neighbors) of a specified metal element only if the number of neighbors equals the given coordination number.

parameters:

coordination_spheres (list): A list of metal coordination spheres (MetalSite objects). metal_element (str): The metal element symbol (e.g., ‘Ag’). coordination_number (int): The desired number of neighbors (coordination number).

Returns:

A list of species symbols (neighbors) bonded to the metal center.

Returns an empty list if no matching environment is found.

Return type:

list[str]

MOF Deconstructions

Deconstruction of frameworks into their chemical building units.

The functions here take an ASE atoms object and work out how the framework is put together: which atoms form a connected component, where the metal to ligand bonds should be cut, and which fragments are symmetry unique. From that the system can be split either into secondary building units, which keep the points of extension, or into metal clusters and whole organic ligands.

Bonding is perceived from geometry using covalent radii, so the input needs sensible coordinates but no bonding information. Cheminformatic identifiers for the resulting fragments come from openbabel when it is installed and from rdkit otherwise, so neither toolkit is a hard requirement; see cheminformatics_backend.

mofstructure.mofdeconstructor.cheminformatics_backend()[source]

Toolkit used to turn a fragment into SMILES, InChI and an InChIKey.

openbabel is preferred because the identifiers shipped with the package, including the IUPAC name database, were built with it. rdkit stands in when openbabel is not installed, so a missing toolkit costs coverage rather than stopping the analysis. Setting MOFSTRUCTURE_CHEMINFO to openbabel or rdkit overrides the choice.

returns:
str

‘openbabel’ or ‘rdkit’.

raises:
ImportError

When neither toolkit is installed, or the requested one is not.

mofstructure.mofdeconstructor.transition_metals()[source]

A function that returns the list of all chemical symbols of metallic elements present in the periodic table

mofstructure.mofdeconstructor.is_metal(symbols)[source]

Check wether symbols in a list are metals

mofstructure.mofdeconstructor.is_alkali(symbols)[source]

Check wether symbols in a list are alkali

mofstructure.mofdeconstructor.inter_atomic_distance_check(ase_atom)[source]

A function that checks whether any pair of atoms (except those with exactly one hydrogen) has a distance below 0.90 Å. Only unique pairs are checked to avoid redundancy.

Parameters

ase_atomase.Atoms

An ASE Atoms object containing atomic positions and chemical symbols.

Returns

bool

Returns False if any applicable atom pair has a distance below 0.90 Å. Otherwise, returns True.

mofstructure.mofdeconstructor.inter_atomic_distance_check2(ase_atom)[source]

Checks whether any non-hydrogen atom in the ASE atoms object has a neighbor within 0.90 Å (ignoring self-distances). This is used to detect overly close contacts, while ignoring R-H bonds (i.e., cases where the central atom is hydrogen).

parameters:
  • ase_atom : ASE atoms object

returns:

bool: False if any non-hydrogen atom has a neighbor (other than itself) within 0.90 Å; True otherwise.

mofstructure.mofdeconstructor.inter_atomic_distance_iterative(ase_atom)[source]

A function that checks whether two atoms are within a distance 1.0 Amstrong unless it is an R-H bond

parameters:
  • ase_atom : ASE atoms object

returns:

boolean :

mofstructure.mofdeconstructor.covalent_radius(element)[source]

A function that returns the covalent radius from the chemical symbol of an element.

parameters:
  • element: str

    Chemical symbol of an atom.

returns:
float

Covalent radius of the element.

mofstructure.mofdeconstructor.ase_2_xyz(atoms)[source]

Create an xyz string from an ase atom object to be compatible with pybel in order to perfom some cheminformatics.

parameters:
  • atoms : ASE atoms object

returns:

a_str : string block of the atom object in xyz format

mofstructure.mofdeconstructor.obmol_2_rdkit(obmol)[source]

A simple function to convert openbabel mol to RDkit mol A function that takes an openbabel molecule object and converts it to an RDkit molecule object. The importance of this function lies in the fact that there is no direct were to convert from an ase atom type to an rdkit molecule object. Consequently, the easiest approach it is firt convert the system to an openbabel molecule object using the function ase_2_pybel to convert to the pybel molecule object. Then using the following code obmol = pybel_mol.OBMol to obtain the openbabel molecule object that can then be used to convert to the rdkit molecule object.

parameters:
  • obmol : openbabel molecule object

returns:

rdmol : rdkit molecule object.

mofstructure.mofdeconstructor.compute_inchis(obmol)[source]

A function to compute the cheminformatic inchi and inchikey using openbabel. The inchi and inchikey are IUPAC cheminformatic identifiers of molecules. They are important to easily search for molecules in databases. For MOFs, this is quite important to search for different secondary building units (sbu) and ligands found in the MOF. More about inchi can be found in the following link

https://iupac.org/who-we-are/divisions/division-details/inchi

parameters:
  • obmol: openbabel molecule object

returns:

inChi, inChiKey: iupac inchi and inchikeys for thesame molecule.

mofstructure.mofdeconstructor.compute_smi(obmol)[source]

A function to compute the SMILES (Simplified molecular-input line-entry system) notation of a molecule using openbabel. The SMILES is a line notation method that uses ASCII strings to represent molecules. More about SMILES can be found in the following link https://en.wikipedia.org/wiki/Simplified_molecular-input_line-entry_system

parameters:
  • obmol : openbabel molecule object

returns:

smi: SMILES notatation of the molecule.

mofstructure.mofdeconstructor.saturate_open_valences(smi)[source]

Fill the open valences left where a ligand was cut away from its metal with implicit hydrogens, recovering the neutral parent molecule. Terephthalate leaves deconstruction as [O]C(=O)c1ccc(cc1)C(=O)[O], two hydrogens short of the terephthalic acid a reference database stores it under.

parameters:
  • smi: SMILES notation of the molecule

returns:

pybel molecule, or None if smi cannot be parsed.

mofstructure.mofdeconstructor.name_lookup_keys(smi)[source]

Identifiers a molecule is indexed under in the IUPAC name database, most specific first: full InChIKey, canonical SMILES, then the InChIKey connectivity block, which ignores protonation. Building the database and querying it both go through here so the two cannot drift apart.

The database was built with openbabel. An InChIKey is defined by the IUPAC algorithm and comes out the same from either toolkit, so the first and third keys still match under rdkit; canonical SMILES is toolkit specific, so the second key only matches the toolkit that wrote the database.

parameters:
  • smi: SMILES notation of the molecule

returns:

list of keys, empty if smi cannot be parsed.

mofstructure.mofdeconstructor.lookup_iupac_name(smi, iupac_names)[source]

Find the IUPAC name of a ligand from the SMILES produced by deconstruction.

parameters:
  • smi: SMILES notation of the ligand, as stored in atoms.info[‘smi’]

  • iupac_names: mapping produced by filetyper.load_iupac_names

returns:

IUPAC name of the ligand, or None when it is not in the database.

mofstructure.mofdeconstructor.ase_2_pybel(atoms)[source]

As simple script to convert from ase atom object to pybel. There are many functionalities like atom typing, bond typing and autmatic addition of hydrogen that can be performed on a pybel molecule object, which can not be directly performed on an ase atom object. E.g

add_hygrogen = pybel.addh()

remove_hydrogen = pybel.removeh()

https://openbabel.org/docs/dev/UseTheLibrary/Python_Pybel.html

parameters:
  • atoms : ASE atoms object

returns:

pybel: pybel molecular object.

mofstructure.mofdeconstructor.max_index(lists)[source]

Extract index of list with the maximum element

mofstructure.mofdeconstructor.compute_openbabel_cheminformatic(ase_atom)[source]

A function the returns all smiles, inchi and inchikeys from an ase_atom object. The function starts by wrapping the system, which is important for systems in which minimum image convertion has made atoms appear to be uncoordinated at random positions. The wrap functions coordinates all the atoms together.

parameters:
  • atoms : ASE atoms object

returns:

smi, inChi, inChiKey

mofstructure.mofdeconstructor.compute_cheminformatic(ase_atom)[source]

SMILES, InChI and InChIKey of a fragment, through whichever toolkit is installed.

openbabel is used when present, since the identifiers distributed with the package were generated with it. rdkit stands in otherwise. The two agree on the InChI and the InChIKey, which are defined by the IUPAC algorithm rather than by the toolkit, but each has its own canonical SMILES, so a SMILES from one will not compare equal to a SMILES from the other.

parameters:
  • ase_atom: ase.Atoms

    The fragment to identify.

returns:
tuple

(smi, inchi, inchikey). Any of them is None when the toolkit could not perceive the fragment.

mofstructure.mofdeconstructor.compute_cheminformatic_from_rdkit(ase_atom)[source]

A function that converts and ase atom into rkdit mol to extract some cheminformatic information.The script begins by taking an ase atom an creating an xyz string of the molecule, which is stored to memory. The xyz string is then read using the MolFromXYZBlock function in rdkit, which reads an xyz string of the molecule. However, this does not include any of the bonding information. To do this the following commands are parsed to the rdkit mol,bond_moll = Chem.Mol(rdmol) rdDetermineBonds.DetermineBonds(bond_moll,charge=0) https://github.com/rdkit/UGM_2022/blob/main/Notebooks/Landrum_WhatsNew.ipynb

parameters:

ASE atoms

returns:

smile string, inchi, inchikey

mofstructure.mofdeconstructor.get_neighbour_bond_matrix(sbu, skin=0.3, bo_step=0.1, aromatic=True)[source]

Create a bond-order connectivity graph using covalent radii and ASE NeighborList, with optional aromatic ring correction.

The algorithm:
  1. Uses covalent radii + tolerance to determine bonded atoms.

  2. Uses progressively shorter cutoffs to heuristically assign single, double, and triple bonds.

  3. Applies MOF-specific rules for metals, alkali metals, and hydrogen atoms.

  4. Optionally detects planar 5-10 membered rings and assigns aromatic bond order (1.5).

parameters:
  • sbuASE atoms object

    The building unit whose connectivity is wanted.

  • skinfloat

    Additional tolerance added to covalent radii (Å). Important for capturing slightly elongated coordination bonds.

  • bo_stepfloat

    Distance reduction step used to define higher bond orders.

  • aromaticbool

    If True, perform aromatic ring detection and correction.

returns:
  1. atom_neighbors:

    A python dictionary where each atom index is a key and the value is a list of bonded atom indices. e.g. atom_neighbors = {0:[1,2,3], 1:[0,4], …}

  2. bonds:

    Bond order matrix (NxN ndarray). Values:

    0.0  -> no bond
    1.0  -> single bond
    2.0  -> double bond (heuristic)
    3.0  -> triple bond (heuristic)
    1.5  -> aromatic bond
    0.5  -> metal-organic coordination-like bond
    0.25 -> metal-metal weak bond
    
mofstructure.mofdeconstructor.compute_ase_neighbour(ase_atom)[source]

Create a connectivity graph using ASE neigbour list.

parameters:

ASE atoms

returns:
  1. atom_neighbors: A python dictionary, wherein each atom index is key and the value are the indices of it neigbours. e.g. atom_neighbors ={0:[1,2,3,4], 1:[3,4,5]…}

  2. matrix: An adjacency matrix that wherein each row correspond to to an atom index and the colums correspond to the interaction between that atom to the other atoms. The entries in the matrix are 1 or 0. 1 implies bonded and 0 implies not bonded.

mofstructure.mofdeconstructor.compute_ase_neighbour_with_offsets(ase_atom)[source]

Create a connectivity graph using ASE neigbour list.

parameters:

ASE atoms

returns:

atom_neighbors: dict[int, list[int]] matrix: adjacency matrix bond_offsets: dict[(i, j)] -> list[(sx, sy, sz)]

mofstructure.mofdeconstructor.matrix2dict(bond_matrix)[source]

A simple procedure to convert an adjacency matrix into a python dictionary.

parameters:

bond matrix : adjacency matrix, type: nxn ndarray

returns:

graph: python dictionary

mofstructure.mofdeconstructor.dfsutil_graph_method(graph, temp, node, visited)[source]

Depth-first search graph algorithm for traversing graph data structures. I starts at the root ‘node’ and explores as far as possible along each branch before backtracking. It is used here a a util for searching connected components in the MOF graph

parameters:
  • graph: any python dictionary

  • temp: a python list to hold nodes that have been visited

  • node: a key in the python dictionary (graph), which is used as the starting or root node

  • visited: python list containing nodes that have been traversed.

returns:

python dictionary

mofstructure.mofdeconstructor.longest_list(lst)[source]

return longest list in list of list

mofstructure.mofdeconstructor.remove_unbound_guest_and_return_unique(ase_atom)[source]

A simple script to remove guest from a metal organic framework. 1)It begins by computing a connected graph component of all the fragments in the system using ASE neighbour list. 2) Secondly it selects indicies of connected components which contain a metal 3)if the there are two or more components, we create a pytmagen graph for each components and filter out all components that are not polymeric 4) If there are two or more polymeric components, we check wether these systems there are identical or different and select only unique polymeric components

parameters:

ASE atoms

returns:

mof_indices : indinces of the guest free system. The guest free ase_atom object can be obtain as follows; E.g. guest_free_system = ase_atom[mof_indices]

mofstructure.mofdeconstructor.remove_unbound_guest(ase_atom)[source]

Remove unbound guest molecules from periodic frameworks and non-periodic molecular systems.

For periodic systems, connected fragments that remain connected after repetition along at least one periodic direction are retained.

For non-periodic systems, the fragment with the greatest total atomic mass is retained. This is useful for organic cages containing smaller solvent or guest molecules.

  1. It begins by computing the connected graph components of all fragments in the system using the ASE neighbour list.

  2. If only one connected component is present, the structure is returned unchanged.

  3. For non-periodic systems (e.g. organic cages), the fragment with the largest total atomic mass is assumed to be the host framework and is retained, while all smaller fragments are treated as guests.

  4. For periodic systems, each connected component is expanded into a supercell along the periodic directions to determine whether it is polymeric.

  5. All polymeric components are retained as the framework, while non-polymeric components are treated as guests.

  6. If no polymeric component can be identified, the fragment with the largest total atomic mass is returned as a fallback.

parameters:
  • ase_atomase.Atoms

    Periodic framework or non-periodic molecular system.

returns:
list[int]

Atom indices belonging to the guest-free host structure.

mofstructure.mofdeconstructor.connected_components_recursive(graph)[source]

Find the connected fragments in a graph. Should work for any graph defined as a dictionary

parameters:

A graph in the form of dictionary e.g.:: graph = {1:[0,1,3], 2:[2,4,5]}

returns:

returns a python list of list of connected components These correspond to individual molecular fragments. list_of_connected_components = [[1,2],[1,3,4]]

mofstructure.mofdeconstructor.dfs_iterative(graph, start, visited)[source]

Depth-first search (DFS) iterative graph algorithm for traversing graph data structures. It starts at the root node ‘start’ and explores as far as possible along each branch before backtracking. This function is used as a utility for identifying connected components in the MOF graph.

parameters:
  • graph: dict

    Keys are nodes, values are lists of neighbouring nodes.

  • start: hashable

    Key in graph used as the root node for the DFS.

  • visited: dict

    Nodes already visited, True when visited and False otherwise.

returns:

A python list containing the nodes in the connected component starting from ‘start’.

mofstructure.mofdeconstructor.connected_components_iterative(graph)[source]

Find the connected fragments in a graph. Should work for any graph defined as a dictionary.

parameters:
  • graph: dict

    Keys are nodes, values are lists of adjacent nodes.

returns:

A python list of lists containing the connected components. Each sublist corresponds to an individual connected component (i.e., a set of nodes that are all connected).

mofstructure.mofdeconstructor.connected_components(graph)[source]

Selects the appropriate connected components algorithm based on the size of the system. Generally, the recursive method is faster but may exceed the recursion depth limit when the system is too large, while the iterative approach is slower but prevents the code from crashing in such cases. This function chooses the connected components method based on the size of the graph.

parameters:
  • graph: dict

    Keys are nodes, values are lists of adjacent nodes.

returns:

A list of lists, where each sublist represents a connected component in the graph.

mofstructure.mofdeconstructor.check_planarity(p1, p2, p3, p4)[source]

A simple procedure to check whether a point is planar to three other points. Important to distinguish porphyrin type metals

parameters:

p1, p2, p3, p4 : ndarray containing x,y,z values

returns:

Boolean True: planar False: noneplanar

mofstructure.mofdeconstructor.metal_in_porphyrin(ase_atom, graph)[source]

A funtion to check whether a metal is found at the centre of a phorphirin. The function specifically identifies metal atoms that are coordinated with four nitrogen atoms, as typically seen in porphyrin complexes. These atoms must also satisfy the planarity condition of the porphyrin structure.

https://en.wikipedia.org/wiki/Transition_metal_porphyrin_complexes

parameters:
  • ase_atom: ASE atom

  • graph: python dictionary containing neigbours

returns:

list of indices consisting of index of metal atoms found in the ASE atom

mofstructure.mofdeconstructor.metal_in_porphyrin2(ase_atom, graph)[source]

https://en.wikipedia.org/wiki/Transition_metal_porphyrin_complexes Check whether a metal is found at the centre of a phorphirin

parameters:
  • ase_atom: ASE atom

  • graph: python dictionary containing neigbours

returns:

list of indices consisting of index of metal atoms found in the ASE atom

mofstructure.mofdeconstructor.move2front(index_value, coords)[source]

Move an index from any position in the list to the front The function is important to set the cell of a rodmof to point in thea-axis. Such that the system can be grow along this axis

parameters:
  • index_value: index of item to move to the from

  • coords: list of coords

returns:

list of coords

mofstructure.mofdeconstructor.find_carboxylates(ase_atom, graph)[source]

A simple aglorimth to search for carboxylates found in the system.

parameters:
  • ase_atom: ASE atom

returns:
  • dictionary of key = carbon index and values = oxygen index

mofstructure.mofdeconstructor.find_nitroxylates(ase_atom, graph)[source]

A simple aglorimth to search for nitroxylates found in the system.

   M
   |
-N-O
   |
   M
parameters:
  • ase_atom: ASE atom

returns:
  • dictionary of key = nitrogen index and values = oxygen index

mofstructure.mofdeconstructor.find_carbonyl_sulphate(ase_atom, graph)[source]

A simple aglorimth to search for Carbonyl sulphate found in the system.

   O
   |
-C-S
   |
   O
parameters:
  • ase_atom: ASE atom

returns:

dictionary of key = carbon index and values = oxygen index

mofstructure.mofdeconstructor.find_sulfides(ase_atom, graph)[source]

A simple aglorimth to search for sulfides.

 S
 |
-C
 |
 S
parameters:
  • ase_atom: ASE atom

returns:

dictionary of key = carbon index and values = sulphur index

mofstructure.mofdeconstructor.find_phosphite(ase_atom, graph)[source]

A simple aglorimth to search for sulfides.

 P
 |
-C
 |
 P
parameters:
  • ase_atom: ASE atom

returns:

dictionary of key = carbon index and values = phosphorous index

mofstructure.mofdeconstructor.find_COS(ase_atom, graph)[source]

A simple aglorimth to search for COS.

O
|
-C
|
S
parameters:
  • ase_atom: ASE atom

returns:

dictionary of key = carbon index and values = sulphur index

mofstructure.mofdeconstructor.find_phosphate(ase_atom, graph)[source]

A simple algorithm to search for Carbonyl sulphate found in the system.

 O
 |
-P-o
 |
 O
parameters:
  • ase_atom: ASE atom

returns:

dictionary of key = carbon index and values = oxygen index

mofstructure.mofdeconstructor.get_bond_shift(i, j, bond_offsets)[source]

Periodic image a bond between two atoms crosses.

A pair can be bonded through more than one image in a small cell, so the most frequent offset is taken as the representative one.

parameters:
  • i, j: atom indices

  • bond_offsets: mapping of atom pair to the cell offsets seen for it

returns:

(x, y, z) cell offset, (0, 0, 0) when the pair is not bonded.

mofstructure.mofdeconstructor.secondary_building_units(ase_atom)[source]
  1. Search for all carboxylate that are connected to a metal. Cut at the position between the carboxylate carbon and the adjecent carbon.

  2. Find all Nitrogen connected to metal. Check whether the nitrogen is in the centre of a porphirin ring. If no, cut at nitrogen metal bond.

  3. Look for oxygen that is connected to metal and two carbon. cut at metal oxygen bond

parameters:
  • ase_atom: ASE atom

returns:

list_of_connected_components: list of connected components,in which each list contains atom indices bonds_to_break: List of lists of atom indices at breaking points porphyrin_checker: Boolean showing whether the metal is in the centre of a porpherin Regions: Dictionary of regions.

mofstructure.mofdeconstructor.rings_and_atom_ring_lookup(graph, use_mcb=True)[source]

Build rings and fast per-atom ring membership lookup.

Parameters:
  • graph (dict[int, list[int]]) – Undirected adjacency list (atom -> neighbors).

  • use_mcb (bool) – If True: use nx.minimum_cycle_basis (SSSR-like). If False: use nx.cycle_basis (fundamental cycles; faster, less “SSSR”).

Returns:

  • rings (list[list[int]]) – List of cycles (each cycle is a list of atom indices).

  • atom_in_ring (dict[int, bool]) – Quick membership: atom_in_ring[i] == True if i is in any ring.

  • atom_to_rings (dict[int, list[int]]) – atom_to_rings[i] gives list of ring indices that contain atom i.

mofstructure.mofdeconstructor.is_atom_in_ring(atom_index, atom_in_ring_lookup)[source]

O(1) check.

mofstructure.mofdeconstructor.ligands_and_metal_clusters(ase_atom)[source]

Start by checking whether there are more than 2 layers if yes, select one Here we select the largest connected component

parameters:
  • ase_atom: ASE atom

returns:

list_of_connected_components: list of connected components, in which each list contains atom indices bonds_to_break: List of lists of atom indices at breaking points Porpyrin_checker: Boolean showing whether the metal is in the centre of a porpherin Regions: Dictionary of regions.

mofstructure.mofdeconstructor.is_rodlike(metal_sbu)[source]

Determine the periodic directions in which a metal SBU remains connected. This helps to identify rod-like SBUs, which are connected in one or two periodic directions.

parameters:
  • metal_sbuase.Atoms

    Metal-containing building unit.

returns:
list[int]

Connected periodic directions: 0 = x, 1 = y, 2 = z. Empty for a non-periodic structure.

mofstructure.mofdeconstructor.all_ferrocene_metals(ase_atom, graph)[source]

A function to find metals corresponding to ferrocene. These metals should not be considered during mof-constructions

parameters:
  • ase_atom: ASE atom

  • graph : dictionary containing neigbour lists

returns:

list_of_metals: list of indices of ferrocene metals

mofstructure.mofdeconstructor.is_ferrocene(metal_sbu, graph)[source]

A simple script to check whether a metal_sbu is ferrocene

parameter:

metal_sbu : ase_atom graph : dictionary containing neigbour lists

returns:

bool : True if the metal_sbu is ferrocene, False otherwise

mofstructure.mofdeconstructor.is_paddlewheel(metal_sbu, graph)[source]

returns True if the atom is part of a paddlewheel motif

parameter:

metal_sbu : ase_atom graph : dictionary containing neigbour lists

returns:

bool : True if the metal_sbu is part of a paddlewheel motif, False otherwise

mofstructure.mofdeconstructor.is_paddlewheel_with_water(ase_atom, graph)[source]

returns True if the atom is part of a paddle wheel with water motif

parameter:

ase_atom : ase_atom graph : dictionary containing neigbour lists

returns:

bool : True if the metal_sbu is part of a paddle wheel with water motif, False otherwise

mofstructure.mofdeconstructor.is_uio66(ase_atom, graph)[source]

returns True if the atom is part of a UIO66 motif

parameter:

ase_atom : ase_atom graph : dictionary containing neigbour lists

returns:

bool : True if the metal_sbu is part of a UIO66 motif, False otherwise

mofstructure.mofdeconstructor.is_irmof(ase_atom, graph)[source]

returns True if the atom is part of a IRMOF motif

parameter:

ase_atom : ase_atom graph : dictionary containing neigbour lists

returns:

bool : True if the metal_sbu is part of a IRMOF motif, False otherwise

mofstructure.mofdeconstructor.is_mof32(ase_atom, graph)[source]

returns True if the atom is part of a MOF32 motif

parameter:

ase_atom : ase_atom graph : dictionary containing neigbour lists

returns:

bool : True if the metal_sbu is part of a MOF32 motif, False otherwise

mofstructure.mofdeconstructor.rod_manipulation(ase_atom, checker)[source]

Script to adjust Rodlike sbus. 1) Its collects the axis responsible for expanding the rod 2) It shifts all coordinates to the axis 3) It rotates the rod to lie in the directions of expansion,

parameter:

ase_atom : ASE atoms object checker : list of indices of atoms responsible for rod expansion

returns:

ase_atom : ASE atoms object with adjusted coordinates

mofstructure.mofdeconstructor.find_unique_building_units(list_of_connected_components, bonds_to_break, ase_atom, porphyrin_checker, all_regions, wrap_system=True, cheminfo=False, add_dummy=False)[source]

A function that identifies and processes unique building units within a MOF. This function deconstructs a MOF into its constituent building unit and identifies unique building units. It can also 1. wrap the system into the unit cell, 2. add dummy atoms to neutralize the building units, 3. compute cheminformatics data using Open Babel.

parameters:
  • list_of_connected_components : list of connected components

  • bonds_to_break : list of lists of atom indices at breaking point

  • ase_atom : ASE atoms object

  • porphyrin_checker : list of indices of porphyrins

  • all_regions : dictionary of regions

  • wrap_system : boolean, default True

  • cheminfo : boolean, default False

  • add_dummy: boolean, default False

    Adds dummy atoms, which can then be replaced by hydrogen to neutralise the building blocks.

returns:

mof_metal: list of metal indices mof_linker: list of linker indices concentration: dictionary of concentration of building units

mofstructure.mofdeconstructor.metal_coordination_number(ase_atom)[source]

Extract coordination number of central metal

paramters:

ase_atom : ASE atoms object

returns:

metal_elt : list of metal elements

mofstructure.mofdeconstructor.metal_coordination_enviroment(ase_atom)[source]

Find the enviroment of a metal

paramters:

ase_atom : ASE atoms object

returns:

metal_elt : list of metal elements metal_enviroment : dict where key is metal element and value is list of neighboring elements

mofstructure.mofdeconstructor.mof_regions(ase_atom, list_of_connected_components, bonds_to_break)[source]

A function to map all atom indices to exact position in which the find themselves in the MOF. This function is used to partition a MOF into regions that correspond to unique unique building units.

parameters:
  • ase_atom: ASE atom

  • list_of_connected_componentslist of list

    Each list holds the atom indices of one building unit.

  • bonds_to_break : list of lists of atom indices at breaking point

returns:

Move an index from any position in the list to the front The function is important to set the cell of a rodmof to point in the a-axis. Such that the system can be grow along this axis

mofstructure.mofdeconstructor.wrap_systems_in_unit_cell(ase_atom, max_iter=30, skin=0.3)[source]

A simple aglorithm to reconnnect all atoms wrapped in a periodic boundary condition such that all atoms outside the box will appear reconnected.

parameters:
  • ase_atom : ASE atoms object

  • max_iter : int, optional (default=30) Maximum number of iterations for reconnection

  • skin : float, optional (default=0.3) Skin distance for bond reconnection

returns:
  • ase_atomASE atoms object

    Reconnected atoms, wrapped in the periodic boundary condition.

mofstructure.mofdeconstructor.angle_tolerance_to_rad(angle, tolerance=5)[source]

Convert angle from degrees to radians with tolerance.

parameters:
  • angle: float

    Angle in degrees.

  • tolerance: float

    Tolerance in degrees, default is 5.

returns:

bond_angle_rad: float Adjusted angle in radians.

mofstructure.mofdeconstructor.find_key_or_value(key_or_value, list_of_list)[source]

Given a value and a list of two-element pairs, return the partner it is paired with. Matching is symmetric: for a pair [i, j] searching i returns j and searching j returns i. This is used to find the atom on the far side of a broken bond, so a dummy can be placed there.

parameters:
  • key_or_valueint

    An atom index to look up.

  • list_of_listlist of pairs

    Broken bonds, each [i, j].

returns:

The partner atom index, or None if the value is not in any pair.

Computing Porosity

Geometric pore analysis through zeo++.

Accessible volume, accessible surface area and the pore diameters (LCD, PLD and the largest free sphere along the percolation path) are computed by handing the structure to zeo++ as a CSSR file.

zeo++ fails in two ways that a plain try/except cannot handle. It aborts the process outright when the Voronoi decomposition fails its internal volume check, which raises SIGABRT rather than a python exception, and on a large or awkward cell it can run for hours without finishing. zeo_calculation therefore runs the calculation in a child interpreter under a timeout, so neither an abort nor a stall can take down a batch job.

Every call returns the same keys whatever happens: a structure that fails gives the same record with None in place of each number and a porosity_status saying why. A directory of structures therefore yields rows of one shape that go straight into a DataFrame, rather than a mixture of records and blanks that has to be repaired before it can be used.

mofstructure.porosity.empty_porosity_record(status)[source]

Build the record returned when zeo++ could not analyse a structure.

The keys match a successful record exactly, so a failure adds a row of missing values to a table rather than a hole in its shape.

parameters:
  • status: str

    Why the calculation produced no numbers.

returns:
python dictionary

Every field in POROSITY_FIELDS set to None, plus porosity_status.

mofstructure.porosity.zeo_calculation(ase_atom, probe_radius=1.86, number_of_steps=10000, high_accuracy=True, rad_file=None, timeout=1800)[source]

Compute the pore geometry of a periodic system in a separate interpreter.

zeo++ validates its Voronoi decomposition against the cell volume and calls abort() when the check fails, which raises SIGABRT rather than a python exception. No try/except can intercept that, so an unlucky structure takes the whole process down with it. A large cell can also leave zeo++ running far longer than the rest of the batch is worth waiting for. Running the calculation in a child under a timeout keeps both to that one structure.

parameters:
  • ase_atom: ase.Atoms

    The periodic structure to analyse.

  • probe_radius: float

    Radius of the probe in Angstrom.

  • number_of_steps: int

    Monte Carlo samples used for the volume and area.

  • high_accuracy: bool

    Use the more expensive Voronoi decomposition.

  • rad_file: str, optional

    File of user defined atomic radii, .rad extension.

  • timeout: float, optional

    Seconds to allow the structure before the child is killed. None waits indefinitely, which risks stalling a batch.

returns:
python dictionary

The fields of POROSITY_FIELDS and porosity_status. A structure zeo++ cannot handle returns the same keys with None in place of each number, so the caller always receives one row of one shape.

mofstructure.porosity.compute_zeo_parameters(ase_atom, probe_radius=1.86, number_of_steps=10000, high_accuracy=True, rad_file=None)[source]

Main script to compute geometric structure of porous systems. The focus here is on MOF, but the script can run on any porous periodic system. The script computes the accesible surface area, accessible volume and the pore geometry. There are many more outputs which can be extracted from ,vol_str and sa_str. Moreover there are also other computation that can be done. Check out the test directory in dependencies/pyzeo/test. Else contact bafgreat@gmail.com. if you need more output and can’t figure it out.

parameter:

ase_atom: ASE atoms object representing the porous structure. probe_radius (float): Radius of the probe (default: 1.86). number_of_steps (int): Number of GCMC simulation cycles (default: 10000). high_accuracy (bool): If True, perform high-accuracy computations. rad_file: Optional file containing user defined atom radii. Must have the .rad extension

returns:

python dictionary containing 1) av_volume_fraction: Accessible volume void fraction 2) av_a3: Accessible volume in A^3 3) asa_a2: Accessible surface area in A^2 4) asa_m2_per_cm3: Accessible surface area in m^2/cm^3 5) number_of_channels: Number of channels present in the porous system, which correspond to the number of pores within the system 6) lcd_a: The largest cavity diameter is the largest sphere that can be inserted in a porous system without overlapping with any of the atoms in the system. 7) pld_a: The pore limiting diameter is the largest sphere that can freely diffuse through the porous network without overlapping with any of the atoms in the system 8) lfpd_a: The largest included sphere along free sphere path is largest sphere that can be inserted in the pore 9) porosity_status: ‘ok’ when the numbers above were computed

Values are plain python floats and ints rather than numpy scalars, so the record serialises to JSON without a custom encoder.

mofstructure.porosity.ase_to_zeoobject(ase_atom)[source]

Converts an ase atom type to a zeo++ Cssr object In zeo++ the xyz coordinate system is rotated to a zyx format.

parameter:

ase_atom: ase atom object

returns:

cssr_object: string representing zeo++ cssr object

Computing stacking in 2D systems and inter-layer height

Stacking analysis of layered covalent organic frameworks.

Layered COFs are described by how consecutive sheets sit relative to one another, which is measured here as the lateral offset between adjacent layers and the interlayer spacing along the stacking direction. That separates eclipsed AA stacking from the various staggered arrangements.

mofstructure.cof_stacking.compute_cof_stacking(ase_atom)[source]

A a simple function to compute the stacking pattern of COFs or layered materials like graphene

parameter:

ase_atom : ASE Atoms object

returns:

layers : list of list wherei each list correspond to a layar lateral_offsets : list of list where each list contains the lateral offsets between two layers interlayer_height : list of list where each list contains the interlayer heights between two layers

mofstructure.cof_stacking.Main()[source]

Main function to compute the stacking pattern of COFs or layered materials like graphene if input is a file it reads directly from the file and if input is a directory it reads all the files in the directory and computes the stacking pattern for each file For directory input it also writes the output to a json file in the same directory

Topological Analysis

Construction of CGD nets from crystal structures.

Topology identification works on nets rather than atoms, so a framework has to be reduced to vertices and edges first. The node definition decides what the net describes, and four are available: sbus places a vertex at each secondary building unit, all_node keeps every branch point (splitting rod SBUs into their atoms), and single_node coarsens all_node by merging each organic group to one vertex. ligand_cluster represents complete organic ligands and metal clusters as the two vertex classes of an incidence net. The same framework can give different nets under each, which is expected rather than an error. cof covers covalent organic frameworks, which have no metal to cut at; see mofstructure.cofstructure.

mofstructure.generate_cgd.remove_broken_bonds_from_breaking_pairs(atom_graph: dict[int, list[int]], bond_offsets: dict[tuple[int, int], list[tuple[int, int, int]]], breaking_pairs: Sequence[Sequence[int | integer]])[source]

Remove all bonds listed in breaking_pairs from an atom-level periodic graph.

The first two entries of each breaking pair correspond to the atom indices of a bond that was cut during deconstruction. This function reconstructs the graph of bonds that remain after those cuts, while preserving the associated lattice offsets of the kept bonds.

parameters:
  • atom_graph: python dictionary

    Atom-level connectivity graph where each key is an atom index and the value is a list of neighbouring atom indices.

  • bond_offsets: python dictionary

    Dictionary mapping (i, j) atom-index tuples to one or more periodic neighbour offsets, e.g. {(i, j): [(sx, sy, sz), …]}.

  • breaking_pairs: list

    List of broken bonds. Each entry is expected to be of the form [i, j] or [i, j, sx, sy, sz].

returns:
kept_graph: python dictionary

Atom-level connectivity graph after removing the broken bonds.

kept_offsets: python dictionary

Dictionary of offsets corresponding to the bonds kept in the graph.

mofstructure.generate_cgd.component_atom_images(components: Sequence[Sequence[int]], kept_graph: dict[int, list[int]], kept_offsets: dict[tuple[int, int], list[tuple[int, int, int]]])[source]

Assign integer lattice image vectors to atoms inside each connected component.

After the bonds listed in breaking_pairs are removed, the structure is split into final connected components. Some of these components may still span the periodic unit cell. This function unwraps each component by traversing the kept internal bonds and assigning an integer image vector (ix, iy, iz) to each atom.

These image vectors are required to compute the true translation between components connected by a broken bond.

parameters:
  • components: list

    List of connected components, where each component is a list of atom indices.

  • kept_graph: python dictionary

    Atom-level connectivity graph after removing broken bonds.

  • kept_offsets: python dictionary

    Dictionary mapping kept bonds to their periodic offsets.

returns:
comp_images: python dictionary
Dictionary of the form:

comp_images[component_id][atom_index] = (ix, iy, iz)

mofstructure.generate_cgd.component_self_translations(components: Sequence[Sequence[int]], kept_graph: dict[int, list[int]], kept_offsets: dict[tuple[int, int], list[tuple[int, int, int]]])[source]

Find the lattice translations along which each component is periodic within itself.

A discrete building unit (a paddlewheel, a Zr6 cluster) is finite: unwrapping its internal bonds assigns every atom one consistent image, so it has no self-translation. A rod-shaped SBU is an infinite chain, so following its internal bonds eventually returns to an atom already seen but in a different periodic image. That image difference is a translation of the rod, and it is exactly the connectivity that keeps a rod framework three-dimensional.

component_atom_images discards these back-edges; this recovers them. The translations found are reduced to a basis of the lattice they generate, so a rod yields one vector, a sheet two, and a discrete cluster none.

parameters:
  • components: list

    Connected components as lists of atom indices.

  • kept_graph: python dictionary

    Atom-level graph after the broken bonds were removed.

  • kept_offsets: python dictionary

    Periodic offsets of the kept bonds.

returns:
python dictionary

Mapping of component id -> basis of the component’s own translation lattice. Empty list for finite components.

mofstructure.generate_cgd.kept_bond_graph(atoms, breaking_pairs)[source]

Build the atom-level graph that remains after the deconstruction cuts.

Wraps the neighbour search and bond removal so a caller that needs the kept graph for more than one purpose (edges between components and each component’s own periodicity, say) can compute it once and pass it around.

parameters:
  • atoms: ASE atoms object

  • breaking_pairs: broken bonds from the deconstructor

returns:

(kept_graph, kept_offsets)

mofstructure.generate_cgd.base_edges_with_shifts(atoms: Atoms, components: Sequence[Sequence[int]], breaking_pairs: Sequence[Sequence[int | integer]], kept_graph: dict[int, list[int]] | None = None, kept_offsets: dict[tuple[int, int], list[tuple[int, int, int]]] | None = None)[source]

Build the base component graph with periodic shifts.

A broken pair returned by mofdeconstructor contains the indices of two atoms that were bonded before deconstruction, and optionally the local bond offset between them. However, this local bond offset alone is not sufficient to determine the true translation between the final post-cut components.

This function therefore:
  1. reconstructs the kept-bond graph,

  2. unwraps each final component using the kept periodic bonds,

  3. computes the translation between components using:

    T(u -> v) = img_u[i] - img_v[j] + s_ij

where:
  • i belongs to component u

  • j belongs to component v

  • img_u[i] is the lattice image of atom i in component u

  • img_v[j] is the lattice image of atom j in component v

  • s_ij is the bond offset stored in breaking_pairs

parameters:
  • atoms: ASE atoms object

  • components: list

    List of final connected components after bond breaking.

  • breaking_pairs: list

    List of broken bonds returned by the deconstructor. Each entry is typically [i, j, sx, sy, sz].

returns:
edges_out: list

List of component-level edges of the form:

(u, v, sx, sy, sz)

where u and v are component indices and (sx, sy, sz) is the corresponding lattice shift.

mofstructure.generate_cgd.component_has_transition_metal_not_in_porphyrin(atoms: Atoms, comp: Sequence[int], porphyrin_atoms: Sequence[int] | None, tm_symbols: Sequence[str])[source]

Check whether a connected component contains a transition metal that is not part of a porphyrin-like motif.

parameters:
  • atoms: ASE atoms object

  • comp: list

    List of atom indices defining a connected component.

  • porphyrin_atoms: list or None

    List of atom indices identified as belonging to a porphyrin motif.

  • tm_symbols: list

    List of transition-metal symbols.

returns:

bool

mofstructure.generate_cgd.regions_with_metal(atoms: Atoms, components: Sequence[Sequence[int]], regions: dict[int, list[int]], porphyrin_atoms: Sequence[int] | None)[source]

Identify regions that contain at least one metal-bearing component.

In this module, a region corresponds to a group of equivalent components returned by the deconstructor. This function selects the regions that contain at least one connected component with a transition metal atom not assigned to a porphyrin motif.

parameters:
  • atoms: ASE atoms object

  • components: list

    List of connected components.

  • regions: python dictionary

    Mapping of region_id -> list of component indices.

  • porphyrin_atoms: list or None

    Atom indices associated with porphyrin-like motifs.

returns:
list

Sorted list of region IDs containing metal-bearing components.

mofstructure.generate_cgd.cgd_from_region_targets(atoms: Atoms, components: Sequence[Sequence[int]], breaking_pairs: Sequence[Sequence[int | integer]], regions: dict[int, list[int]], *, target_regions: Sequence[int], name: str = 'net', dedup_edges: bool = True)[source]

Build a CGD PERIODIC_GRAPH by contracting non-target components into edges.

The logic is:
  1. Build the base component graph with periodic shifts.

  2. Select all components belonging to target_regions as CGD nodes.

  3. Treat all remaining components as edge-components.

  4. Contract each edge-component into effective edges between node-components.

The contraction is incidence-based, meaning that different periodic incidences between a linker-like component and node-like components are preserved. This is essential for recovering the correct 3D net of pillared or otherwise periodic frameworks.

parameters:
  • atoms: ASE atoms object

  • components: list

    List of connected components from the deconstructor.

  • breaking_pairs: list

    List of broken atom pairs from the deconstructor.

  • regions: python dictionary

    Mapping of region_id -> list of component indices.

  • target_regions: list

    Region IDs that should be treated as node regions.

  • name: str

    CGD graph ID.

  • dedup_edges: bool

    If True, remove exact duplicate periodic edges.

returns:
cgd_text: str

CGD PERIODIC_GRAPH content as a string.

mofstructure.generate_cgd.dedup_periodic_edges(edges)[source]

Remove duplicate undirected periodic edges. An edge (u, v, s) and its reverse (v, u, -s) are the same edge; a self-edge (u, u, s) and (u, u, -s) likewise.

mofstructure.generate_cgd.periodic_graph_cgd(edges, name)[source]

Format a list of (u, v, sx, sy, sz) edges (0-based nodes) as a CGD PERIODIC_GRAPH that Systre can read.

mofstructure.generate_cgd.coordination_contacts(atoms: Atoms, components: Sequence[Sequence[int]], breaking_pairs: Sequence[Sequence[int | integer]], atom_graph: dict[int, list[int]], is_cluster: Sequence[bool])[source]

Select the cut bonds that are genuine ligand-to-cluster coordination.

ligands_and_metal_clusters reports a breaking pair for every rule that fires, and some of those pairs name two atoms that are not bonded to each other: the carboxylate rule, for instance, records the carboxylate carbon against the metal its oxygen is bound to. Roughly half the pairs returned for a carboxylate framework are of that kind, and they carry no real bond offset, so an edge built from one would sit at an arbitrary image. Only pairs that are bonds in the atom graph, and that join a ligand component to a cluster component, describe a contact.

parameters:
  • atoms: ASE atoms object

  • components: connected components from the deconstructor

  • breaking_pairs: broken bonds from the deconstructor

  • atom_graph: atom-level connectivity of the intact structure

  • is_cluster: per component, True when the component is a metal cluster

returns:

contacts: list of (metal_atom, donor_atom, cluster_id, ligand_id)

atom_component: mapping atom index -> component id

mofstructure.generate_cgd.ligand_cluster_incidences(atoms: Atoms)[source]

Deconstruct a framework into ligand and cluster vertices and find every coordination incidence between them.

The vertices are the components of ligands_and_metal_clusters, which cuts the structure at the coordination bonds themselves: a ligand keeps every atom it owns, including its donor atoms, and a cluster keeps its metals with whatever inorganic bridges hold them together, so every atom belongs to exactly one vertex. An incidence is a contact between a ligand and one periodic image of a cluster, and it carries its denticity – the number of donor bonds making that one contact – because a carboxylate that has lost one oxygen to a defect still makes the same contact but with a lower denticity.

A component that is periodic within itself, a rod or a sheet, is one vertex however far along its own chain a contact is made, so contact translations are reduced modulo that component’s own translation lattice. Without this, chemically equivalent linkers of a rod framework come out with different coordination numbers and the answer changes when the same crystal is given in a supercell.

parameters:
  • atoms: ASE atoms object, guest-free

returns:
python dictionary with keys:

components: the vertex components, as atom-index lists is_cluster: per component, True for a metal cluster lattices: per component, basis of its own translation lattice incidences: mapping (cluster, ligand, sx, sy, sz) -> denticity

mofstructure.generate_cgd.ligand_cluster_graph(atoms: Atoms, *, collapse_ditopic: bool = False)[source]

Construct the periodic incidence net between ligands and metal clusters.

Vertices and edges come from ligand_cluster_incidences: every complete ligand and every metal cluster is a vertex, and an edge records that a ligand coordinates one periodic image of a cluster. Several donor bonds within the same contact (a chelating or bridging carboxylate) are one edge, so chelation does not inflate the topological degree.

A ligand held by a single contact – coordinated solvent, a capping modulator, or a linker left dangling by a missing-linker defect – is a vertex of degree one. No periodic net can carry one: it collides in the barycentric placement, so no canonical key exists and those ligands are left out of the graph. The defect is not lost, it shows in the reduced connectivity of the cluster that lost the linker, and ligand_cluster_fingerprint records the dangling ligand itself.

parameters:
  • atoms: ASE atoms object, guest-free

  • collapse_ditopic: bool

    If True, a ligand bridging exactly two cluster contacts is spliced into a direct cluster–cluster edge instead of staying a vertex. The framework is unchanged, but the net is no longer subdivided, so it can carry an RCSR name (UiO-66 reads fcu rather than a 12,2-net that has no entry in the archive). Left False, the ligand stays a vertex and the net describes how each ligand meets the clusters.

returns:

edges: list of periodic graph edges

ligand_nodes: set of node identifiers belonging to ligands

node_atoms: mapping from node identifiers to represented atom indices

mofstructure.generate_cgd.splice_two_connected_nodes(edges, candidates)[source]

Replace each two-connected candidate vertex by a direct edge.

A ditopic ligand is a connection rather than a branch point, and keeping it as a vertex subdivides the edge it makes. The subdivided net is a faithful description but has no RCSR name – the archive holds pcu, not its subdivision – so splicing restores the net that can be identified.

parameters:
  • edges: list of (u, v, sx, sy, sz)

  • candidates: node ids allowed to be spliced out

returns:

edges: list of edges with the spliced vertices removed

removed: set of node ids that were spliced out

mofstructure.generate_cgd.ZEOLITE_T_ELEMENTS = frozenset({'Al', 'As', 'B', 'Be', 'Co', 'Fe', 'Ga', 'Ge', 'Li', 'Mg', 'P', 'Si', 'Ti', 'V', 'Zn'})

Elements that sit at the tetrahedral sites of a zeolite framework. The list follows the IZA framework-type compositions and is deliberately permissive at the transition-metal end, since Zn, Co and Fe zeolite analogues are all known.

mofstructure.generate_cgd.zeolite_t_edges(atoms: Atoms, *, bridge: Sequence[str] = ('O',), t_elements: Sequence[str] | None = None)[source]

Contract a zeolite framework onto its tetrahedral atoms.

A zeolite has no metal cluster to cut at, so none of the deconstruction methods apply to it. Its net is instead the classical one of the crystallographic literature: vertices are the tetrahedral (T) atoms and an edge is a T-O-T bridge, the oxygen contracted away. Delgado-Friedrichs & O’Keeffe use exactly this net in their worked example, where “the edges of the net correspond to -O- links between Si atoms at the vertices of the net”.

The contraction is not cosmetic. Keeping the bridging oxygen as a vertex describes the subdivision of the net, in which every edge is split in two, and that is a different graph: it keys perfectly well and matches nothing, because the archive names the contracted net. The same is true of any pendant atom left attached, which quietly decorates the net rather than raising, so hydrogen and other terminal species are dropped here rather than being left to produce a silent miss.

Shifts compose the way the path does. If a bridge sees its two T neighbours in cells s_i and s_j, the contracted edge carries s_j - s_i, which is the translation from the first T atom to the second.

parameters:
  • atoms: ASE atoms object

    Framework, guests already removed.

  • bridge: sequence of str

    Elements that are contracted away, two-coordinate by assumption.

  • t_elements: sequence of str, optional

    Elements accepted at a tetrahedral site. Defaults to ZEOLITE_T_ELEMENTS. Anything neither bridge nor T - an extra-framework cation, say - takes no part in the net.

returns:
tuple

(n_vertices, edges, report) with edges as (u, v, sx, sy, sz) over 0-based T indices, and report recording the T atoms kept and any bridge that did not join exactly two of them.

mofstructure.generate_cgd.cgd_zeolite(atoms: Atoms, *, name: str = 'net', bridge: Sequence[str] = ('O',), t_elements: Sequence[str] | None = None)[source]

Build the CGD periodic graph of a zeolite framework from its T atoms.

parameters:
  • atoms: ASE atoms object

    Framework, guests already removed.

  • name: str

    CGD graph ID.

  • bridge: sequence of str

    Elements contracted away.

  • t_elements: sequence of str, optional

    Elements accepted at a tetrahedral site.

returns:
str

CGD PERIODIC_GRAPH content.

raises:
RuntimeError

If no tetrahedral atoms are present, or none of them are bridged, which means the input is not a framework of this kind.

mofstructure.generate_cgd.cgd_ligand_cluster(atoms: Atoms, *, name: str = 'net', collapse_ditopic: bool = False)[source]

Return a CGD representation of the ligand–metal-cluster incidence net.

mofstructure.generate_cgd.sbu_drawing_graph(atoms: Atoms)[source]

Construct the SBU-contracted graph together with its atom mapping.

mofstructure.generate_cgd.ligand_cluster_fingerprint(atoms: Atoms, *, algorithm: str = 'sha256')[source]

Describe how ligands meet metal clusters.

Net identification answers a narrower question than this one. It names the net when the net has a name, but a ligand–cluster net keeps ditopic ligands as vertices and so is a subdivided net that RCSR does not list, and a framework with a dangling ligand has no stable net at all. This fingerprint is computed from the deconstruction itself, so it exists for every framework, and it is invariant to atom order, to the choice of cell origin, and to being given a supercell of the same crystal.

What it records is chemistry that a net alone cannot carry. Each ligand and cluster species is counted by formula, with a histogram of how many clusters each ligand bridges and of the denticity of those contacts. A missing-linker defect shows as a cluster species that has lost contacts; a linker left hanging by one end shows under terminal with its own formula, which is what separates it from a coordinated solvent; a carboxylate reduced from bridging to monodentate shows in the denticity histogram while the net itself is unchanged.

Every count is quoted per metal-cluster repeat unit, as an exact fraction: pristine HKUST-1 reads 4/3 of a C9H3O6 ligand per Cu2 paddlewheel, and one coordinated methanol in that cell reads 1/24.

parameters:
  • atoms: ASE atoms object

    Guest-free framework.

  • algorithm: str

    Hash algorithm supported by hashlib.

returns:
python dictionary

With keys:

  • clusters: species -> {count, contacts histogram, capped histogram}

  • ligands: species -> {count, contacts histogram, denticity histogram}

  • terminal: species -> {count, denticity histogram}, for ligands held by a single contact

  • refinement: histogram of refined colours, as a sorted list of (colour digest, count)

  • cluster_units: metal-cluster repeat units in the cell, the denominator the counts are quoted against

  • fingerprint_hash: hex digest of everything above

mofstructure.generate_cgd.cgd_all_nodes(atoms: Atoms, components: Sequence[Sequence[int]], breaking_pairs: Sequence[Sequence[int | integer]], regions: dict[int, list[int]], *, target_regions: Sequence[int], name: str = 'net')[source]

Build the all-node CGD, where a rod SBU keeps its atoms as separate nodes.

A discrete metal cluster collapses to one node, exactly as cgd_from_region_targets does, so the net for a discrete-SBU framework is unchanged. A rod SBU is an infinite chain, so collapsing it to one node throws away the chain and drops the framework to a lower-dimensional net. It is instead split: each metal atom and each bridging carboxyl carbon becomes a node, and the oxygen atoms between them contract to edges. This recovers the true net – MIL-53 gives rna rather than the collapsed pcu.

parameters:
  • atoms: ASE atoms object

  • components: connected components from the deconstructor

  • breaking_pairs: broken bonds from the deconstructor

  • regions: mapping region_id -> component ids

  • target_regions: region ids treated as node (metal) regions

  • name: CGD graph ID

returns:

cgd_text: str

mofstructure.generate_cgd.cgd_single_nodes(atoms: Atoms, components: Sequence[Sequence[int]], breaking_pairs: Sequence[Sequence[int | integer]], regions: dict[int, list[int]], *, target_regions: Sequence[int], name: str = 'net')[source]

Build the single-node CGD, a coarsening of the all-node net.

Starting from the all-node net, every connected group of organic vertices (the carboxyl-carbon and linker nodes) is merged into a single vertex, while the metal vertices are left as they are. A rod therefore keeps its metals as separate nodes but collapses each linker to one vertex. This is the representation CrystalNets calls SingleNodes; MIL-53 gives bpq. An organic group that is itself periodic is left un-merged, so it cannot swallow a whole periodic chain.

Arguments are the same as cgd_all_nodes.

mofstructure.generate_cgd.net_geometry(atoms, *, method='all_node')[source]

Trace the topological net over the real structure geometry.

Each net node is placed at the real (Cartesian) centroid of the atoms it represents, so the returned net follows the actual framework rather than the idealised RCSR cell. Intended for drawing, not for topology identification.

parameters:
  • atoms: guest-free ASE atoms object

  • method: “sbus”, “all_node”, “single_node” or “ligand_cluster”

returns:

positions: dict node_id -> (x, y, z) Cartesian centroid kinds: dict node_id -> “metal” or “organic” edges: list of (u, v, sx, sy, sz) with periodic shifts cell: 3x3 cell matrix (rows are lattice vectors)

mofstructure.generate_cgd.drawing_periodic_edges(positions, edges, cell)[source]

Express periodic edges in the coordinate gauge used by drawn node positions.

Periodic graph translations depend on the unit-cell representative chosen for every node. Node centroids are wrapped independently for display, so their representatives can differ from those used during graph construction. Integer gauge shifts are propagated along a spanning forest to reconcile the two choices without changing cycle translations. This preserves periodic self-edges and distinct connections between different images of the same pair of nodes.

class mofstructure.generate_cgd.TopologyExtractor(ase_atoms: Atoms | None = None, filename: str | None = None)[source]

Bases: object

A high-level class for extracting a periodic topology from a porous crystal.

The class accepts either an ASE atoms object or a structure filename readable by ASE. It removes unbound guests, performs deconstruction using one of the supported methods, selects node regions, contracts linker-like regions, and writes the resulting CGD PERIODIC_GRAPH.

parameters:
  • ase_atoms: ASE atoms object, optional

  • filename: str, optional

    Path to an input structure file readable by ASE.

methods:
remove_guest():

Remove unbound guests from the framework.

build_cgd():

Build a CGD periodic graph using one of the supported topology modes.

write_cgd():

Write CGD content to file.

ase_atoms: Atoms | None = None
filename: str | None = None
property atoms

The structure being analysed, read from filename when needed.

returns:

ase.Atoms

remove_guest()[source]

Remove unbound guests from a porous periodic structure.

returns:
ASE atoms object

Guest-free framework.

build_cgd(*, method: str = 'sbus', name: str = 'net')[source]

Build a CGD PERIODIC_GRAPH using one of the supported methods.

parameters:
  • method: str
    One of:
    • “sbus”: each SBU is one node (rods collapse with a self-edge).

    • “all_node”: rod SBUs keep their atoms as separate nodes, recovering the true net (e.g. MIL-53 gives rna, not pcu).

    • “single_node”: the all-node net with each organic group merged to one vertex (CrystalNets SingleNodes; MIL-53 gives bpq).

    • “ligand_cluster”: complete organic ligands and metal clusters form the two vertex classes of an incidence net.

    • “cof”: for a covalent organic framework, which has no metal to cut at. The cut is made at the linkage bond instead, and a building unit is a vertex when it is joined at three or more points.

    • “zeol”: for a zeolite, which likewise has nothing to cut at. Vertices are the tetrahedral atoms and each T-O-T bridge becomes an edge, the oxygen contracted away, which is the net the zeolite literature names (ABW gives sra, EDI gives edi).

  • name: str

    CGD graph ID.

returns:
cgd_text: str

CGD PERIODIC_GRAPH content.

static write_cgd(cgd_text, path)[source]

Write CGD content to file.

parameters:
  • cgd_text: str

    CGD content.

  • path: str

    Output file path.

returns:
str

Output path.

mofstructure.generate_cgd.build_argparser()[source]

Build the command-line argument parser.

returns:

argparse.ArgumentParser

mofstructure.generate_cgd.main(argv: Sequence[str] | None = None)[source]

Command-line entry point.

parameters:
  • argv: list, optional

    Arguments; sys.argv[1:] when omitted.

returns:
int

0 when a net was written, 2 when the framework could not be reduced to one.

Topology Identification

One call from a structure to its topology.

Everything needed to identify the net of a framework already exists in this package - generate_cgd deconstructs a structure into a quotient graph and graph_net identifies one - but nothing joined them, so a caller had to deconstruct, write a CGD, parse it back, build a PeriodicGraph and call identify by hand, and had to know which deconstruction a given material wanted before starting. This module is that missing layer.

one answer shape for every material. A MOF, a COF and a zeolite are deconstructed along completely different lines - at the metal cluster, at the linkage bond, at the bridging oxygen - but what comes back is the same record in all three cases, because the thing being reported is a net and a net does not remember what it was made of. That is what makes the results comparable: a MOF and a COF built on the same net answer with the same key_hash, which they would not if each material class had its own reply.

what is always there, and what is not. key is the identification. It exists for any locally stable net that finishes computing, and it is unique: nets with the same key are the same net, and 17409 independent nets were checked to give 17409 distinct keys. A name is a different matter, and depends on whether an archive happens to carry the net. Three situations produce no key at all, and each is reported as itself rather than as a failure: a net that is unstable, for which no canonical form exists at all; a fragment with no periodicity, which is not a net; and a net too large to finish inside the budget.

references:
Delgado-Friedrichs, O. & O’Keeffe, M. (2003). Acta Cryst. A59, 351-360.

doi:10.1107/S0108767303012017

O’Keeffe, M., Peskov, M. A., Ramsden, S. J. & Yaghi, O. M. (2008).

Acc. Chem. Res. 41, 1782-1789. (RCSR)

Ramsden, S. J., Robins, V. & Hyde, S. T. (2009). Acta Cryst. A65, 81-108.

(EPINET)

mofstructure.topology.DEFAULT_METHOD = {'cof': 'cof', 'mof': 'all_node', 'zeolite': 'zeol'}

The deconstruction each kind of framework is identified with by default. A COF is cut at its linkage bond and a zeolite at its bridging oxygen, and in neither case is there a second sensible place to cut.

mofstructure.topology.MOF_METHODS = ('sbus', 'all_node', 'single_node', 'ligand_cluster')

The alternative nets a MOF has, in preference order. No one net is more true than another; they answer different questions. sbus contracts each metal cluster to a vertex, all_node keeps the atoms of a rod apart, so MIL-53 reads as rna rather than pcu, and ligand_cluster gives an incidence net sensitive to defects.

mofstructure.topology.is_cgd(structure: str | Atoms) bool[source]

Whether the input is a CGD periodic graph rather than a structure.

Accepts a path ending in .cgd and CGD text passed directly, so a net written by this package can be read straight back without going through a structure file it never came from.

parameters:
  • structure: str or ASE atoms object

    Candidate input.

returns:

bool

mofstructure.topology.analyse(structure: str | Atoms, *, method: str = 'auto', timeout: int | None = 300, descriptors: bool = False, symmetry: bool = False, name: str = 'net') dict[str, object][source]

Identify the topology of a framework.

parameters:
  • structure: str or ASE atoms object

    Path to anything ASE can read, or the atoms themselves.

  • method: str

    Deconstruction to use. “auto” classifies the structure first; otherwise any method TopologyExtractor.build_cgd accepts.

  • timeout: int or None

    Seconds allowed for the identification, None for no limit. The cost of the canonical form is driven by the number of vertices and by degenerate coordination rather than by the size of the structure, and a large net can take hours, so a bounded answer that says it ran out of time is more useful than an unbounded one that never returns.

  • descriptors: bool

    Compute coordination sequences and point and vertex symbols. These are the useful output for a net with no name, and the most expensive part of the call.

  • symmetry: bool

    Derive the maximal space group and intrinsic chirality.

  • name: str

    Name to carry through to the quotient graph.

returns:
python dictionary

Keys status, material, method, key_version, components and, when every component agrees, the convenience fields topology, key and key_hash. status is “ok” when a canonical key was produced, and otherwise names what stopped it: “deconstruction_failed” when the framework could not be reduced to building units, “no_net” when the deconstruction left no periodic graph, “unidentified” when a graph was built but has no canonical form, which happens for an unstable net, and “timeout” or “error”. A named net is not required for “ok”. topology reads as one of three things and never as nothing: a name, UNNAMED_TOPOLOGY for a net that was identified but that no archive names, and FAILED_TOPOLOGY for a structure that produced no net, whatever status says stopped it. Which archive supplies a name depends on the material, following NAME_PREFERENCE_BY_MATERIAL. The per-component records are those graph_net.identify returns.

mofstructure.topology.analyse_methods(structure: str | Atoms, methods: Sequence[str] | None = None, **options) dict[str, dict[str, object]][source]

Identify one structure under every deconstruction that suits it.

This is a MOF question. A MOF has more than one defensible net and they answer different questions, so where a single answer is wanted analyse picks the conventional one and this reports them all side by side. A COF and a zeolite have one cut each and answer with that alone. The structure is read once and reused, since reading dominates the cost for small nets.

parameters:
  • structure: str or ASE atoms object

    Structure to identify.

  • methods: sequence of str, optional

    Deconstructions to run. Defaults to MOF_METHODS for a MOF and to the single sensible method for a COF or a zeolite.

  • options:

    Passed through to analyse.

returns:
python dictionary

Mapping method name -> the record analyse returns for it. A method that does not apply is present with its failing status rather than omitted, so the absence of an answer is visible.

mofstructure.topology.analyse_many(structures: Sequence[str | Atoms], **options) list[dict[str, object]][source]

Identify a sequence of structures, one record each.

A structure that fails is recorded and the run continues, since the common use is a directory of thousands where a handful will not deconstruct.

parameters:
  • structures: sequence

    Paths or atoms objects.

  • options:

    Passed through to analyse.

returns:
list

One record per structure, in input order, each carrying source.

mofstructure.topology.as_atoms(structure: str | Atoms) Atoms[source]

Accept either a structure file or an ASE atoms object.

parameters:
  • structure: str or ASE atoms object

    Path to anything ASE can read, or the atoms themselves.

returns:

ASE atoms object

mofstructure.topology.classify(structure: str | Atoms) str[source]

Decide which kind of framework a structure is.

The order of the tests matters. A zeolite is looked for first and is identified by what it lacks: no carbon, and nothing outside the tetrahedral elements and their bridging oxygen. Several of those tetrahedral elements are metals, so asking “does it contain a metal” first would call every zeolite a MOF. Once a zeolite is ruled out, a metal means a MOF and its absence means a COF.

parameters:
  • structure: str or ASE atoms object

    Structure to classify.

returns:
str

One of “zeolite”, “mof” or “cof”.

mofstructure.topology.quotient_graph(cgd_text: str, name: str = 'net') PeriodicGraph[source]

Read a CGD PERIODIC_GRAPH block back into a quotient graph.

The CGD that generate_cgd writes is an explicit list of (tail, head, shift) edges rather than a geometric description, so this conversion loses nothing and no coordinates are involved.

parameters:
  • cgd_text: str

    CGD content holding a PERIODIC_GRAPH block.

  • name: str

    Name to give the resulting graph.

returns:

PeriodicGraph

raises:
ValueError

If the text holds no EDGES block, which is what an empty deconstruction produces.

Periodic Net Analysis

Pure-Python identification and analysis of periodic nets.

graph_net replaces the parts of the Systre pipeline that required a Java runtime. Given the quotient graph of a net it computes a canonical key, looks that key up against the RCSR archive, derives the maximal symmetry the net can achieve, and writes a geometric realisation as CGD. Nothing in the path uses the unit cell, the space group or any relaxed geometry of the structure the net came from: identification is combinatorial, and symmetry and geometry are derived from the topology rather than supplied to it.

the pieces.

periodic_graph the quotient graph, lattice reduction, components primitive reduction to the primitive cell barycentric exact equilibrium placement and stability canonical the canonical key archive RCSR lookup invariants coordination sequences, TD10, point and vertex symbols symmetry maximal space group and intrinsic chirality embedding geometric realisation and CGD output lattice exact integer lattice arithmetic underneath all of it

on keys. The canonical form is not byte-compatible with Systre’s. The published algorithm leaves several tie-breaks open and the conventions in the shipped archive differ from the printed ones, so rather than reverse-engineer an unspecified implementation this package uses its own canonical form and re-keys the RCSR archive with it. What is guaranteed, and tested, is the property that matters: one key for every representation of a net, and different keys for different nets. A key produced here must not be compared against a key produced by Systre.

references:
Delgado-Friedrichs, O. & O’Keeffe, M. (2003). Acta Cryst. A59, 351-360.

doi:10.1107/S0108767303012017

Delgado-Friedrichs, O. (2005). Discrete Comput. Geom. 33, 67-81.

doi:10.1007/s00454-004-1147-x

Delgado-Friedrichs, O. & O’Keeffe, M. (2005). J. Solid State Chem. 178,

2480-2485. doi:10.1016/j.jssc.2005.06.011

O’Keeffe, M., Peskov, M. A., Ramsden, S. J. & Yaghi, O. M. (2008).

Acc. Chem. Res. 41, 1782-1789. doi:10.1021/ar800124u

class mofstructure.graph_net.PeriodicGraph(dim: int, n_vertices: int, edges: tuple[tuple[int, int, tuple[int, ...]], ...], name: str | None = None)[source]

Bases: object

A finite encoding of a periodic net.

parameters:
  • dim: int

    Number of columns of every shift vector, i.e. the rank of the translation group used to write the representation down. This is an upper bound on the periodicity, not the periodicity itself; use periodicity() for that.

  • n_vertices: int

    Number of vertex orbits.

  • edges: tuple

    Tuple of normalised (u, v, shift) triples, sorted and deduplicated.

  • name: str or None

    Optional label carried through from the input file.

dim: int
n_vertices: int
edges: tuple[tuple[int, int, tuple[int, ...]], ...]
name: str | None = None
classmethod build(dim: int, n_vertices: int, edges: Iterable[Sequence], name: str | None = None) PeriodicGraph[source]

Construct a graph, normalising and deduplicating the edges.

A net is a simple graph, so two identical (u, v, s) triples describe one edge and the duplicate is dropped. A self-loop with zero shift joins a vertex to itself in the same unit cell and is not an edge of any net, so it is rejected.

parameters:
  • dim: int

    Ambient translation rank.

  • n_vertices: int

    Number of vertex orbits.

  • edges: iterable

    Iterable of (u, v, shift) with zero-based vertex indices.

  • name: str or None

    Optional label.

returns:

PeriodicGraph

property n_edges: int

Number of edge orbits.

returns:

int

incidences() dict[int, list[tuple[int, tuple[int, ...]]]][source]

Map each vertex orbit to the images of its neighbours.

For a vertex v the list holds one (w, s) pair per incident edge end, so that the neighbour sits at the image of w translated by s. A self-loop contributes two entries, +s and -s, which is what makes the degree of a looped vertex come out as two per loop.

returns:
python dictionary

Mapping vertex index -> list of (neighbour, shift).

degrees() list[int][source]

Coordination number of each vertex orbit.

Chemists call this the coordination number rather than the degree or valency (Delgado-Friedrichs & O’Keeffe 2005, section 2).

returns:
list

One integer per vertex orbit.

is_quotient_connected() bool[source]

Whether the finite quotient graph is connected.

This is weaker than connectivity of the infinite net: the quotient can be connected while the cover falls apart into translated copies.

returns:

bool

spanning_tree() tuple[dict[int, tuple[int, ...]], list[tuple[int, ...]]][source]

Breadth-first spanning tree potentials and the resulting cycle vectors.

Each vertex v of a connected quotient graph is assigned the accumulated translation q(v) reached along tree edges from vertex 0, with q(0) = 0. Every edge (u, v, s) that is not a tree edge then closes a cycle whose net translation is

c = q(u) + s - q(v).

The set of all such c generates the translation lattice of the net. This is the standard cycle-space construction behind the genus of a quotient graph (Delgado-Friedrichs & O’Keeffe 2005, section 5).

returns:
tuple

(potentials, cycle_vectors) where potentials maps vertex index to its accumulated translation.

translation_lattice() list[tuple[int, ...]][source]

Canonical basis of the lattice of translations realised by the net.

returns:
list

Rows of the Hermite basis; its length is the periodicity.

periodicity() int[source]

Number of independent directions in which the net actually repeats.

A net written in a 3-dimensional cell whose cycle lattice has rank 2 is a 2-periodic net, and Delgado-Friedrichs & O’Keeffe are explicit that such a net must not be called 3-dimensional (2005, section 2).

returns:

int

covering_multiplicity() int | None[source]

Number of translated copies the infinite cover splits into.

When the cycle lattice L has full rank, the vertices of the cover fall into [Z^dim : L] classes that are never joined by an edge, so the net drawn in the given cell is that many identical interpenetrating copies. An index of 1 means the cover is connected.

returns:
int or None

The index, or None when the cycle lattice is rank deficient.

reduced() PeriodicGraph[source]

Rewrite the shifts in the net’s own translation lattice.

This performs two normalisations at once. A rank-deficient cycle lattice is projected onto its own rank, so a 2-periodic net presented in a 3-dimensional cell comes back with two-column shifts. A full-rank sublattice of index k > 1 is rescaled to index 1, which is what makes the representation independent of the cell the caller happened to supply and is the reason a 1x1x2 supercell cannot change the answer.

returns:
PeriodicGraph

Equivalent graph whose cycle lattice is the full integer lattice of its own dimension.

components() list[PeriodicGraph][source]

Split into the connected components of the quotient graph.

An interpenetrated or catenated framework deconstructs into several disjoint nets, and each has to be identified separately; Systre reports them as numbered components. Vertices are renumbered from zero inside each component.

returns:
list

One PeriodicGraph per component, ordered by smallest original vertex index.

genus() int[source]

Cyclomatic number of the quotient graph.

For a connected quotient graph with v vertex orbits and e edge orbits

genus = 1 + e - v,

which Beukemann & Klee showed is at least the periodicity n, with equality defining the minimal nets (Delgado-Friedrichs & O’Keeffe 2005, section 5).

returns:

int

key_string() str[source]

Serialise as the whitespace-separated integer string used by Systre.

The layout is the dimension followed by, for each edge, the two one-based vertex indices and the shift components. This is the format of the key records in the RCSR archive; it is only a canonical key when the graph itself is already in canonical form.

returns:

str

classmethod from_key_string(text: str, name: str | None = None) PeriodicGraph[source]

Parse the integer string produced by key_string.

parameters:
  • text: str

    Dimension followed by (u, v, shift…) records.

  • name: str or None

    Optional label.

returns:

PeriodicGraph

relabelled(permutation: Sequence[int]) PeriodicGraph[source]

Apply a permutation of the vertex orbits.

parameters:
  • permutation: sequence

    permutation[i] is the new index of old vertex i.

returns:

PeriodicGraph

transformed(matrix: Sequence[Sequence[int]]) PeriodicGraph[source]

Apply a unimodular change of lattice basis to every shift.

The new shift is matrix @ shift. A unimodular matrix leaves the net unchanged and only rewrites the representation, so this is the natural way to generate equivalent inputs when testing a canonical form.

parameters:
  • matrix: sequence

    Square integer matrix of side dim, determinant +/-1.

returns:

PeriodicGraph

translated(offsets: dict[int, Sequence[int]]) PeriodicGraph[source]

Move individual orbit representatives to other images of themselves.

Choosing a different representative for vertex orbit v shifts every edge end at v, without changing the net. Delgado-Friedrichs & O’Keeffe point this out as one of the sources of the infinitely many representations a single net admits (2003, section 2).

parameters:
  • offsets: python dictionary

    Mapping vertex index -> integer translation.

returns:

PeriodicGraph

exception mofstructure.graph_net.UnstableNetError(message: str, collisions: list | None = None)[source]

Bases: ValueError

Raised when a net is not locally stable and cannot be canonicalised.

parameters:
  • message: str

    Human readable description.

  • collisions: list

    List of colliding items, either vertex pairs or (vertex, position) groups depending on which check failed.

mofstructure.graph_net.canonical_form(graph: PeriodicGraph, check_stability: bool = True, return_winners: bool = False)[source]

Canonical representative of the isomorphism class of graph.

The graph is first reduced onto its own translation lattice and then onto its primitive cell, so that neither the choice of cell nor a rank-deficient or index-greater-than-one lattice can influence the answer. Both steps are needed and they are different: PeriodicGraph.reduced normalises the lattice, while primitive.primitive detects vertex orbits related by a translation, which is how an ordinary 1x1x2 supercell presents itself.

parameters:
  • graph: PeriodicGraph

    Quotient graph with a connected quotient. Use PeriodicGraph.components() first if it may be disconnected.

  • check_stability: bool

    When True, refuse to canonicalise a net that is not locally stable rather than returning a key that depends on tie-breaking.

  • return_winners: bool

    When True, also return the reduced graph and every traversal that attained the canonical form. Two traversals reaching the same canonical representation differ by an automorphism of the net, so this list is the raw material of the symmetry computation (Delgado-Friedrichs, GD 2003, LNCS 2912, section 5: “whenever two d-tuples of directed edges lead to the same representation, an automorphism of the periodic graph has been found and all automorphisms must occur in this way”).

returns:
PeriodicGraph

Canonical representative; its key_string() is the canonical key. When return_winners is set, a triple of the canonical graph, the reduced graph and the list of winning traversals.

raises:
UnstableNetError

If the net is not locally stable and check_stability is set.

ValueError

If the quotient graph is disconnected or has no valid basis.

mofstructure.graph_net.canonical_key(graph: PeriodicGraph, check_stability: bool = True) str[source]

Canonical key of a net.

parameters:
  • graph: PeriodicGraph

    Quotient graph with a connected quotient.

  • check_stability: bool

    Refuse unstable nets when True.

returns:
str

Whitespace-separated integer string.

mofstructure.graph_net.describe(graph: PeriodicGraph, shells: int = 10, max_size: int = 12) dict[str, object][source]

Full metric-free description of a net.

This is what should be reported when the canonical key is absent from the RCSR archive: the net has no name, but it is completely characterised for a reader by these invariants together with its key.

parameters:
  • graph: PeriodicGraph

    Net to describe.

  • shells: int

    Shells for the coordination sequence.

  • max_size: int

    Cap for cycle and ring searches.

returns:
python dictionary

Coordination numbers, sequences, TD, point and vertex symbols, genus, periodicity and interpenetration multiplicity.

mofstructure.graph_net.ideal_embedding(graph: PeriodicGraph, scale: str = 'mean') dict[str, object][source]

Symmetric geometric realisation of a net.

parameters:
  • graph: PeriodicGraph

    Net with a connected quotient graph.

  • scale: str

    mean normalises the mean edge length to one, min the shortest.

returns:
python dictionary

Keys lattice (three vectors as rows), cell (the six parameters), positions (fractional, one per vertex orbit), edges (the quotient edges), edge_lengths, and edge_length_spread, the ratio of longest to shortest edge, which is the honest measure of how far this embedding is from one with uniform edges.

raises:
ValueError

If the net is not 3-periodic, since the cell construction and the CGD format both assume three dimensions.

mofstructure.graph_net.describe_key(key: str, prefer: Sequence[str] | None = None) dict[str, str | None]

Everything the archive can say about a canonical key.

The name and its source are reported separately because they are different claims. An RCSR symbol says the net is one reticular chemistry has named; an EPINET symbol says it matches an entry in a systematic enumeration of nets obtained by reticulating hyperbolic tilings. Both are useful, and conflating them would overstate the second.

A net with no name at all is an ordinary outcome, not a failure: the key still identifies it uniquely, which is what makes two structures comparable to each other whether or not anyone has named their topology.

One net may be named by more than one archive: a zeolite framework carries an RCSR symbol and an IZA framework-type code at once, so ABW is sra to the former and ABW to the latter. topology therefore reports a single headline name and topology_source says which archive it came from, with names carrying every name known for the net. That shape keeps one field to read for “what is this”, states the provenance explicitly, and takes a further archive without changing its structure - where one field per archive would need a new field and every caller would have to learn about it.

Preference runs RCSR, then IZA, then EPINET: a name reticular chemistry has chosen says more than a framework code, which in turn says more than a position in a systematic enumeration.

parameters:
  • key: str

    Canonical key as produced by canonical.canonical_key.

  • prefer: sequence of str, optional

    Archives the headline name may be drawn from, most preferred first. Defaults to NAME_PREFERENCE. A caller that knows what the material is can narrow it, so that a MOF is never headlined with a zeolite framework code; names still carries every name, so narrowing the headline discards nothing.

returns:
python dictionary

Keys key, key_hash, key_version, topology, topology_source and names. The last three are None, None and an empty dictionary for a net no archive has named, which is an ordinary outcome rather than a failure.

mofstructure.graph_net.identify(graph: PeriodicGraph, descriptors: bool = True, symmetry: bool = True, prefer: Sequence[str] | None = None) list[dict[str, object]][source]

Identify every component of a net and describe it.

A framework may deconstruct into several disjoint nets, as an interpenetrated structure does, so each connected component is handled separately and the result is a list. A component that cannot be keyed is reported with its reason rather than dropped, because silently omitting a component is how an interpenetrated structure gets mistaken for a simple one.

parameters:
  • graph: PeriodicGraph

    Net to identify; it may be disconnected.

  • descriptors: bool

    Compute coordination sequences and point and vertex symbols. These are the useful output for a net with no RCSR name, and the most expensive part of the call.

  • symmetry: bool

    Derive the maximal space group and intrinsic chirality.

  • prefer: sequence of str, optional

    Archives the headline topology may come from, most preferred first; passed straight through to archive.describe.

returns:
list

One dictionary per component, each carrying periodicity, n_vertices, n_edges, interpenetration, the canonical key with its key_hash and key_version, the headline name topology with the topology_source it came from and every known name in names, all of which are empty for a net no archive has named. When requested, descriptors and symmetry follow. A component that could not be keyed carries error instead of a key, which happens when it is unstable or has no periodicity at all.

mofstructure.graph_net.key_hash(key: str, algorithm: str = 'sha256') str[source]

Stable fixed-width handle for a canonical key.

The key itself is the identity of a net and should be stored, since only it can be looked up again if the archive later grows. A hash of it is convenient as a database column or join field, and it is a hash of the key rather than of a CGD file for a concrete reason: a CGD describes one representation, so the same net written with a different vertex order, origin or supercell hashes differently, while its canonical key - and therefore this digest - does not.

parameters:
  • key: str

    Canonical key as produced by canonical.canonical_key.

  • algorithm: str

    Any name accepted by hashlib.new.

returns:
str

<version>:<algorithm>:<hexdigest>, the version included so a digest made under a different canonical form is never mistaken for one made under this one.

mofstructure.graph_net.lookup(key: str) str | None[source]

RCSR identifier for a canonical key, or None when the net is not named.

A miss is a perfectly ordinary outcome and means the net is new to the archive rather than that anything failed; Systre reports the same situation as “structure is new for this run”.

parameters:
  • key: str

    Canonical key as produced by canonical.canonical_key.

returns:

str or None

mofstructure.graph_net.primitive(graph: PeriodicGraph) PeriodicGraph[source]

Rewrite a net on its primitive cell.

Detects the additional translations, extends the lattice by them and collapses each orbit of vertices onto a single representative, re-expressing every edge through equation (3). A net already given primitively is returned unchanged.

parameters:
  • graph: PeriodicGraph

    Net whose quotient graph is connected.

returns:
PeriodicGraph

Equivalent net with the smallest number of vertex orbits and the finest translation lattice.

mofstructure.graph_net.space_group(graph: PeriodicGraph, symprec: float = 1e-05) dict[str, object][source]

Maximal crystallographic symmetry of a net, named by spglib.

parameters:
  • graph: PeriodicGraph

    Net with a connected quotient graph.

  • symprec: float

    Tolerance handed to spglib. The operations are exact integers, so this only guards the floating point lattice built from the metric.

returns:
python dictionary

Keys order, international, number, hall, crystal_system, is_chiral and rotations. Naming keys are None for a net that is not 3-periodic, since the 230 space group types only classify three dimensions.

mofstructure.graph_net.to_cgd(graph: PeriodicGraph, name: str | None = None, key: str | None = None, rcsr: str | None = None, refine: bool = False) str[source]

Write a net as a CGD CRYSTAL block.

The block is emitted in space group P1 with every vertex orbit listed explicitly. That is a deliberate choice over quoting the ideal space group and an asymmetric unit: the orbit representatives produced here are in the net’s own primitive setting, which need not be one of the conventional settings the symbol implies, and a CGD file whose coordinates and group do not agree is worse than one with no symmetry claimed at all. The ideal group is recorded in a comment instead, where it cannot be misread as a generator instruction.

parameters:
  • graph: PeriodicGraph

    Net to write.

  • name: str or None

    Name for the NAME record.

  • key: str or None

    Canonical key, written as a comment so an unnamed net stays identifiable from its own file.

  • rcsr: str or None

    Archive name when the net is a named one: an RCSR symbol for a MOF or COF, an IZA framework-type code for a zeolite.

returns:
str

CGD text.