R-group isomers¶
Counting and enumeration of isomers with two or more different functional
groups. Used by CageBuilder.count_rgroup_isomers() and
CageBuilder.enumerate_rgroup_isomers(): cage_isomer_builder.utils.rgroup.
Symmetry-unique isomers with several distinct functional groups placed at once (R1, R2, ..., written here as letters A, B, C, ...).
Two ways of specifying what goes on the linkers:
groups(default mode): every linker carries the same set of groups.groups="AB"puts exactly one A and one B on every linker,groups="AAB"two A and one B,groups="A"reproduces the classic one-group-per-linker isomers.linker_patterns(mixed mode): linkers carry different sets, with a fixed number of linkers of each kind:{"A": 3, "B": 3}puts one A on three linkers and one B on the other three,{"AB": 2, "": 4}leaves four linkers bare. The values must sum to the number of linkers.
Mathematical basis
Let X be the FG anchor slots, partitioned into linkers of F slots each,
and G the symmetry group acting on X as permutations. A configuration is a
labelling f: X -> {0 (H), 1 (A), 2 (B), ...}. G acts by
(g.f)(g(x)) = f(x), and two configurations are the same isomer iff they lie
in the same G-orbit. Burnside's lemma:
n_unique = (1/|G|) * sum_{g in G} |Fix(g)|
f is fixed by g iff f is constant on every cycle of g. Every symmetry maps
whole linkers onto whole linkers (checked, see :func:prepare_group), so g
also permutes the linkers, and a fixed f must
give every linker in one linker cycle the same set of groups. A site cycle
inside a linker cycle of length l visits each of its l linkers the same
number of times w = |cycle| / l. Hence
|Fix(g)| = [ prod_P y_P^{N_P} ] prod_{linker cycles Lam}
( sum_P W(Lam, P) * y_P^{l(Lam)} )
where P runs over the allowed linker patterns, N_P is how many linkers must carry P, and W(Lam, P) is the number of ways to give each site cycle of Lam one label so that each linker receives exactly the composition of P:
W(Lam, P) = #{ labels for the cycles : sum of w over cycles labelled c
= n_c(P) for every group c }.
Both coefficients are extracted with small exact dynamic programmes over Python integers (no overflow, no floating point). For the identity this reduces to the multinomial count
|Fix(e)| = (n_linkers! / prod_P N_P!) * prod_P ( F! / ((F-|P|)! prod_c n_c(P)!) )^{N_P}
Two assumptions are built in. Groups with different letters are never interchanged by symmetry (A and B are chemically different), and the substituents are achiral (improper operations such as mirrors count as symmetries, exactly as for the single-group counts).
Enumeration keeps a configuration iff it is the lexicographically smallest member of its orbit, compared first on the per-linker pattern sequence and then on the per-slot labels. Because of that ordering, a whole linker-level pattern assignment can be skipped when it is not itself minimal under the induced linker permutations. A complete enumeration is always checked against the Burnside count.
SlotLayout
dataclass
¶
How the FG slots are grouped into linkers.
Slots are numbered linker by linker: linker L owns the
sizes[L] consecutive slots starting at offsets[L]. Linkers of
different chemical types (e.g. BDC and BPDC in one MOF) may have
different numbers of slots; types[L] is the type index of linker
L and type_names[t] its name. A symmetry operation may only send
a linker onto a linker of the same type.
Attributes:
| Name | Type | Description |
|---|---|---|
sizes |
tuple of int
|
|
types |
tuple of int
|
|
type_names |
tuple of str
|
|
Examples:
>>> from cage_isomer_builder.utils.rgroup import SlotLayout
>>> layout = SlotLayout(sizes=(4, 4, 2), types=(0, 0, 1), type_names=("bdc", "pyz"))
>>> layout.n_slots, layout.offsets, layout.slot_linker.tolist()
(10, (0, 4, 8), [0, 0, 0, 0, 1, 1, 1, 1, 2, 2])
n_linkers
property
¶
n_slots
property
¶
offsets
property
¶
slot_linker
property
¶
uniform_size
property
¶
The common number of slots per linker, or None if they differ.
Examples:
n_types
property
¶
uniform
classmethod
¶
n_linkers linkers of one type with fg_per_linker slots.
Examples:
>>> from cage_isomer_builder.utils.rgroup import SlotLayout
>>> SlotLayout.uniform(6, 4).sizes
(4, 4, 4, 4, 4, 4)
Source code in cage_isomer_builder/utils/rgroup.py
linkers_of_type ¶
Indices of the linkers of type t.
Examples:
>>> from cage_isomer_builder.utils.rgroup import SlotLayout
>>> layout = SlotLayout(sizes=(4, 4, 2), types=(0, 0, 1), type_names=("bdc", "pyz"))
>>> layout.linkers_of_type(0), layout.linkers_of_type(1)
([0, 1], [2])
Source code in cage_isomer_builder/utils/rgroup.py
type_size ¶
Number of slots of each linker of type t (ValueError if they differ).
Examples:
>>> from cage_isomer_builder.utils.rgroup import SlotLayout
>>> layout = SlotLayout(sizes=(4, 4, 2), types=(0, 0, 1), type_names=("bdc", "pyz"))
>>> layout.type_size(0), layout.type_size(1)
(4, 2)
Source code in cage_isomer_builder/utils/rgroup.py
RGroupSpec
dataclass
¶
RGroupSpec(labels: tuple, patterns: tuple, pattern_names: tuple, pattern_counts: tuple, pattern_types: tuple, layout: SlotLayout)
Normalised description of what goes on the linkers.
Attributes:
| Name | Type | Description |
|---|---|---|
labels |
tuple of str
|
Group names; label integer |
patterns |
tuple of tuple of int
|
Per-pattern composition: |
pattern_names |
tuple of str
|
|
pattern_counts |
tuple of int
|
Number of linkers that must carry each pattern. |
pattern_types |
tuple of int
|
Linker type each pattern belongs to. |
layout |
SlotLayout
|
|
Examples:
>>> from cage_isomer_builder.utils.rgroup import make_spec
>>> spec = make_spec(6, 4, linker_patterns={"A": 3, "B": 3})
>>> spec.labels, spec.patterns, spec.pattern_counts
(('A', 'B'), ((1, 0), (0, 1)), (3, 3))
SlotGroup
dataclass
¶
SlotGroup(perms: ndarray, inverse: ndarray, linker_perms: ndarray, linker_inverse: ndarray, layout: SlotLayout = None)
A validated symmetry group acting on FG slots.
Attributes:
| Name | Type | Description |
|---|---|---|
perms |
np.ndarray, shape (|G|, n_slots)
|
|
inverse |
np.ndarray, shape (|G|, n_slots)
|
|
linker_perms |
np.ndarray, shape (|G|, n_linkers)
|
Induced permutation of whole linkers. |
linker_inverse |
np.ndarray, shape (|G|, n_linkers)
|
|
layout |
SlotLayout
|
|
Examples:
>>> from cage_isomer_builder.utils import rgroup
>>> # two linkers with two slots each (slots 0,1 | 2,3); one symmetry swaps the linkers
>>> group = rgroup.prepare_group([[2, 3, 0, 1]], n_slots=4, fg_per_linker=2)
>>> group.order, group.linker_perms.tolist()
(2, [[0, 1], [1, 0]])
order
property
¶
RGroupIsomer
dataclass
¶
One symmetry-unique isomer.
Attributes:
| Name | Type | Description |
|---|---|---|
site_labels |
tuple of int
|
Label per FG slot: 0 = H, i = |
labels |
tuple of str
|
|
linker_patterns |
tuple of str
|
Pattern carried by each linker. |
Examples:
>>> from cage_isomer_builder.utils.rgroup import RGroupIsomer
>>> iso = RGroupIsomer(site_labels=(1, 0, 2, 0), labels=("A", "B"), linker_patterns=("A", "B"))
>>> iso.to_dict(), iso.name
({'A': [0], 'B': [2]}, 'A-0_B-2')
name
property
¶
slots ¶
FG slot indices carrying group label (e.g. "A").
Examples:
>>> from cage_isomer_builder.utils.rgroup import RGroupIsomer
>>> RGroupIsomer((1, 1, 0, 2), ("A", "B"), ("AA", "B")).slots("A")
[0, 1]
Source code in cage_isomer_builder/utils/rgroup.py
to_dict ¶
as_layout ¶
A :class:SlotLayout from a layout or a uniform fg_per_linker.
Examples:
>>> from cage_isomer_builder.utils.rgroup import as_layout
>>> as_layout(4, n_slots=24).sizes # six linkers of four slots
(4, 4, 4, 4, 4, 4)
Source code in cage_isomer_builder/utils/rgroup.py
make_spec ¶
Build an :class:RGroupSpec.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
layout
|
SlotLayout or int
|
The slot layout, or (single linker type) the number of linkers, with
|
required |
groups
|
str, list, or dict
|
Groups carried by every linker (default mode). With several linker
types, a dict |
'AB'
|
linker_patterns
|
dict
|
Mixed mode (takes precedence): |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If a pattern needs more slots than its linker has, two patterns of one type are the same set of groups written differently, or the linker counts of a type don't add up to its number of linkers. |
Examples:
>>> from cage_isomer_builder.utils.rgroup import SlotLayout
>>> layout = SlotLayout(sizes=(4, 4, 2), types=(0, 0, 1), type_names=("bdc", "pyz"))
>>> from cage_isomer_builder.utils.rgroup import make_spec
>>> make_spec(6, 4, groups="AAB").patterns # two A and one B on every linker
((2, 1),)
>>> spec = make_spec(layout, groups={"bdc": "AB", "pyz": ""}) # per linker type
>>> spec.pattern_names, spec.pattern_types
(('AB', ''), (0, 1))
Source code in cage_isomer_builder/utils/rgroup.py
365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 | |
prepare_group ¶
Turn a transformation library into a validated :class:SlotGroup.
Takes the operation-major layout (a list of permutations, as taken by
count_unique_isomers). The slot-major output of
CageBuilder._get_transformation_library must be transposed first
with list(zip(*tl)); the two layouts cannot be told apart when the
number of operations equals the number of slots. The identity
is added and duplicates removed.
Every condition Burnside's lemma relies on is checked rather than assumed: each operation is a bijection of the slots, maps whole linkers onto whole linkers of the same type, and the set is closed under composition.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fg_per_linker
|
int or SlotLayout
|
Slots per linker, or the layout of linkers with different numbers of slots (several linker types). |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If any of those checks fails. This always means the symmetry operations don't match the structure (e.g. a mis-oriented cage or a too-loose tolerance), and any count derived from them would be wrong. |
Examples:
>>> from cage_isomer_builder.utils import rgroup
>>> # two linkers with two slots each (slots 0,1 | 2,3); one symmetry swaps the linkers
>>> group = rgroup.prepare_group([[2, 3, 0, 1]], n_slots=4, fg_per_linker=2)
>>> group.perms.tolist() # identity added
[[0, 1, 2, 3], [2, 3, 0, 1]]
>>> rgroup.prepare_group([[1, 2, 3, 0]], 4, 2) # moves slot 1 onto the other linker
Traceback (most recent call last):
...
ValueError: a symmetry operation splits the slots of one linker across several linkers; per-linker compositions are then not symmetry invariant.
Source code in cage_isomer_builder/utils/rgroup.py
511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 | |
identity_fixed_count ¶
|Fix(e)|: the raw number of configurations before symmetry.
Examples:
>>> from cage_isomer_builder.utils.rgroup import identity_fixed_count, make_spec
>>> identity_fixed_count(make_spec(6, 4, groups="A")) # 4**6 raw placements
4096
>>> identity_fixed_count(make_spec(2, 4, groups="AB")) # (4 * 3)**2
144
Source code in cage_isomer_builder/utils/rgroup.py
count_rgroup_isomers ¶
count_rgroup_isomers(transformation_library, n_slots, fg_per_linker=4, groups='AB', linker_patterns=None)
Exact number of symmetry-unique isomers with several distinct groups, via Burnside's lemma (see the module docstring for the formula).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
transformation_library
|
list or SlotGroup
|
Symmetry operations on the FG slots, operation-major (see
:func: |
required |
n_slots
|
int
|
Total number of FG slots. |
required |
fg_per_linker
|
int or SlotLayout
|
Slots per linker, or the layout for several linker types. |
4
|
groups
|
str, sequence of str, or dict
|
Groups carried by every linker (default mode); per linker type with
a dict (see :func: |
"AB"
|
linker_patterns
|
dict
|
Mixed mode (see :func: |
None
|
Returns:
| Type | Description |
|---|---|
int
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
Invalid specification, invalid symmetry group (see
:func: |
Examples:
>>> from cage_isomer_builder.utils import rgroup
>>> # two linkers with two slots each (slots 0,1 | 2,3); one symmetry swaps the linkers
>>> group = rgroup.prepare_group([[2, 3, 0, 1]], n_slots=4, fg_per_linker=2)
>>> rgroup.count_rgroup_isomers(group, 4, 2, groups="A") # (2*2 raw + 2 fixed) / 2
3
>>> rgroup.count_rgroup_isomers(group, 4, 2, linker_patterns={"A": 1, "": 1})
2
Source code in cage_isomer_builder/utils/rgroup.py
iter_rgroup_isomers ¶
iter_rgroup_isomers(transformation_library, n_slots, fg_per_linker=4, groups='AB', linker_patterns=None, limit=None, batch_size=4096)
Generator over symmetry-unique isomers with several distinct groups.
One representative per orbit: the lexicographically smallest configuration, compared first on the pattern sequence over the linkers and then on the slot labels. Memory use is independent of the size of the configuration space. The inner loop is compiled with numba.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
transformation_library
|
As for :func: |
required | |
n_slots
|
As for :func: |
required | |
fg_per_linker
|
As for :func: |
required | |
groups
|
As for :func: |
required | |
linker_patterns
|
As for :func: |
required | |
limit
|
int
|
Stop after this many isomers. |
None
|
batch_size
|
int
|
How many isomers the compiled loop collects per call. |
4096
|
Yields:
| Type | Description |
|---|---|
RGroupIsomer
|
|
Examples:
>>> from cage_isomer_builder.utils import rgroup
>>> # two linkers with two slots each (slots 0,1 | 2,3); one symmetry swaps the linkers
>>> group = rgroup.prepare_group([[2, 3, 0, 1]], n_slots=4, fg_per_linker=2)
>>> # one representative per orbit: (0, 1, 1, 0) also stands for (1, 0, 0, 1)
>>> [iso.site_labels for iso in rgroup.iter_rgroup_isomers(group, 4, 2, groups="A")]
[(0, 1, 0, 1), (0, 1, 1, 0), (1, 0, 1, 0)]
Source code in cage_isomer_builder/utils/rgroup.py
976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 | |
enumerate_rgroup_isomers ¶
enumerate_rgroup_isomers(transformation_library, n_slots, fg_per_linker=4, groups='AB', linker_patterns=None, limit=None)
List of symmetry-unique isomers (see :func:iter_rgroup_isomers).
A complete enumeration (limit=None) is checked against the
independent Burnside count; a mismatch raises RuntimeError.
Examples:
>>> from cage_isomer_builder.utils import rgroup
>>> # two linkers with two slots each (slots 0,1 | 2,3); one symmetry swaps the linkers
>>> group = rgroup.prepare_group([[2, 3, 0, 1]], n_slots=4, fg_per_linker=2)
>>> isomers = rgroup.enumerate_rgroup_isomers(group, 4, 2, groups="AB")
>>> len(isomers), isomers[0].to_dict()
(3, {'A': [0, 2], 'B': [1, 3]})
Source code in cage_isomer_builder/utils/rgroup.py
build_rgroup_isomer_atoms ¶
The isomer as ase.Atoms. Every active slot is an R site, an X
atom whose per-atom "rgroup" entry is its group number (1 = R1 =
first label, 2 = R2, ...; 0 for every other atom), see
:mod:cage_isomer_builder.utils.sites. Inactive slots go back to H at
their original position (or 1.09 Angstrom along the C-site direction if
that is unknown). atoms.info["rgroup_labels"] records the group
names, e.g. "A B".
Examples:
>>> import numpy as np
>>> from ase.build import molecule
>>> from cage_isomer_builder.utils.sites import mark_site
>>> benzene = molecule("C6H6") # C 0-5, H 6-11 (H i+6 on C i)
>>> h_positions = benzene.positions[6:].copy()
>>> for h in range(6, 12):
... mark_site(benzene, h, 1) # every H a slot, as functionalise() does
>>> slots = list(range(6, 12))
>>> anchors = benzene[slots]
>>> from cage_isomer_builder.utils.rgroup import RGroupIsomer, build_rgroup_isomer_atoms
>>> iso = RGroupIsomer((1, 0, 0, 2, 0, 0), ("NH2", "OH"), ("NH2", "OH"))
>>> atoms = build_rgroup_isomer_atoms(benzene, anchors, iso, slots, h_positions)
>>> atoms.get_chemical_symbols()[6:], atoms.info["rgroup_labels"]
(['X', 'H', 'H', 'X', 'H', 'H'], 'NH2 OH')
Source code in cage_isomer_builder/utils/rgroup.py
write_rgroup_isomer_file ¶
write_rgroup_isomer_file(cage, fg_anchors, isomer, fg_anchor_indices, output_path, fg_anchor_h_positions=None)
Write the isomer to <output_path>/<isomer.name>.xyz in extended XYZ
format, which keeps the rgroup labels and, for MOFs, the unit cell.
Read back with ase.io.read; atoms.arrays["rgroup"] has the labels.
Returns:
| Type | Description |
|---|---|
str
|
Path of the written file. |
Examples:
>>> import numpy as np
>>> from ase.build import molecule
>>> from cage_isomer_builder.utils.sites import mark_site
>>> benzene = molecule("C6H6") # C 0-5, H 6-11 (H i+6 on C i)
>>> h_positions = benzene.positions[6:].copy()
>>> for h in range(6, 12):
... mark_site(benzene, h, 1) # every H a slot, as functionalise() does
>>> slots = list(range(6, 12))
>>> anchors = benzene[slots]
>>> import tempfile
>>> from cage_isomer_builder.utils.rgroup import RGroupIsomer, write_rgroup_isomer_file
>>> from cage_isomer_builder.utils.sites import read_isomer
>>> iso = RGroupIsomer((1, 0, 0, 2, 0, 0), ("NH2", "OH"), ("NH2", "OH"))
>>> path = write_rgroup_isomer_file(benzene, anchors, iso, slots, tempfile.mkdtemp(), h_positions)
>>> path.endswith("NH2-0_OH-3.xyz"), read_isomer(path).info["r_sites"]
(True, {6: 'R1', 9: 'R2'})
Source code in cage_isomer_builder/utils/rgroup.py
rgroup_atoms_to_rdkit ¶
RDKit molecule for a finite isomer, with every active site as an RDKit
R-group dummy atom (R# in a MOL/SDF file, [1*:1] in SMILES), the
R1, R2, ... notation RDKit uses for R-group decomposition.
Bonds come from bond_matrix (e.g. CageBuilder.get_bond_matrix),
not from distance perception, which is unreliable on metal clusters.
The molecule is not sanitised. Periodic structures should stay in
extended XYZ, since SDF cannot store a unit cell.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
atoms
|
Atoms
|
Output of :func: |
required |
bond_matrix
|
(ndarray, shape(len(atoms), len(atoms)))
|
Bond orders. 1.5 is written as an aromatic bond; other orders outside 1-4 (e.g. fractional metal-ligand codes) as single bonds. |
required |
Examples:
>>> import numpy as np
>>> from ase.build import molecule
>>> from cage_isomer_builder.utils.sites import mark_site
>>> benzene = molecule("C6H6") # C 0-5, H 6-11 (H i+6 on C i)
>>> h_positions = benzene.positions[6:].copy()
>>> for h in range(6, 12):
... mark_site(benzene, h, 1) # every H a slot, as functionalise() does
>>> slots = list(range(6, 12))
>>> anchors = benzene[slots]
>>> from rdkit import Chem
>>> from cage_isomer_builder.utils.rgroup import RGroupIsomer, build_rgroup_isomer_atoms, rgroup_atoms_to_rdkit
>>> iso = RGroupIsomer((1, 0, 0, 2, 0, 0), ("NH2", "OH"), ("NH2", "OH"))
>>> atoms = build_rgroup_isomer_atoms(benzene, anchors, iso, slots, h_positions)
>>> from cage_isomer_builder.utils.features import bond_graph
>>> bonds = np.zeros((12, 12))
>>> for i, nbrs in bond_graph(atoms).items():
... for j in nbrs:
... bonds[i, j] = 1.5 if max(i, j) < 6 else 1.0
>>> block = Chem.MolToMolBlock(rgroup_atoms_to_rdkit(atoms, bonds))
>>> [line for line in block.splitlines() if line.startswith("M RGP")]
['M RGP 2 7 1 10 2']
Source code in cage_isomer_builder/utils/rgroup.py
1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 | |
rgroup_smiles ¶
CXSMILES of a finite isomer in which every site is an R-group atom.
Each site is written as a dummy atom * carrying an atom label, so the
string ends in e.g. |$;;_R1;;_R2$|, the CXSMILES atom labels marking
R1 and R2. The same molecule written to SDF/MOL (via
:func:rgroup_atoms_to_rdkit and RDKit's MolToMolBlock) shows
R# atoms with M RGP records.
Examples:
>>> import numpy as np
>>> from ase.build import molecule
>>> from cage_isomer_builder.utils.sites import mark_site
>>> benzene = molecule("C6H6") # C 0-5, H 6-11 (H i+6 on C i)
>>> h_positions = benzene.positions[6:].copy()
>>> for h in range(6, 12):
... mark_site(benzene, h, 1) # every H a slot, as functionalise() does
>>> slots = list(range(6, 12))
>>> anchors = benzene[slots]
>>> from cage_isomer_builder.utils.rgroup import RGroupIsomer, build_rgroup_isomer_atoms, rgroup_smiles
>>> iso = RGroupIsomer((1, 0, 0, 2, 0, 0), ("NH2", "OH"), ("NH2", "OH"))
>>> atoms = build_rgroup_isomer_atoms(benzene, anchors, iso, slots, h_positions)
>>> from cage_isomer_builder.utils.features import bond_graph
>>> bonds = np.zeros((12, 12))
>>> for i, nbrs in bond_graph(atoms).items():
... for j in nbrs:
... bonds[i, j] = 1.5 if max(i, j) < 6 else 1.0
>>> smiles = rgroup_smiles(atoms, bonds)
>>> smiles.split(" ")[0]
'[H]c1c([H])c([2*:2])c([H])c([H])c1[1*:1]'
>>> "_R1" in smiles and "_R2" in smiles
True