Interactive Quantum Chemistry
git clone git@github.com:Autonomous-Scientific-Agents/IQC.git
cd IQCuv is a fast Python package installer and resolver written in Rust.
-
Install uv (if not already installed):
curl -LsSf https://astral.sh/uv/install.sh | shOr using pip:
pip install uv
-
Create a virtual environment and install IQC:
uv venv source .venv/bin/activate # On Windows: .venv\Scripts\activate uv pip install -e .
Or in a single command:
uv pip install -e . --python 3.12
-
First, install a package manager (Conda, Miniconda, Mamba, or MicroMamba)
- Download Miniconda from the official page
- Follow the installation instructions for your operating system.
-
Create and activate the environment:
conda env create -f env.yml conda activate iqc-env
-
Install IQC:
pip install .
The base IQC install does not install PyTorch, MACE, UMA, or e3nn. Install
the mlip extra only when you need MLIP calculators.
On Linux, installing PyTorch dependencies from the default PyPI index can pull CUDA/NVIDIA wheels even on systems without NVIDIA GPUs. For CPU-only or Intel GPU systems, install the appropriate PyTorch wheel first, then install IQC's MLIP runtime dependencies.
CPU-only PyTorch:
uv pip install torch --index-url https://download.pytorch.org/whl/cpu
uv pip install -e ".[mlip]"
uv pip install --no-deps mace-torchIntel GPU / XPU PyTorch:
uv pip install torch --index-url https://download.pytorch.org/whl/xpu
uv pip install -e ".[mlip]"
uv pip install --no-deps mace-torchFor NVIDIA GPU systems, install the CUDA PyTorch wheel appropriate for the
machine, then run the same .[mlip] and mace-torch --no-deps steps.
When installing MACE and FAIRChem UMA in the same Python 3.11+ environment,
install a hardware-appropriate PyTorch build first, install IQC's MLIP runtime
dependencies, then install mace-torch without dependencies so its old
e3nn==0.4.4 metadata does not downgrade the newer UMA-compatible e3nn
stack. If PyTorch is already installed correctly for your machine, the MLIP
steps are:
uv pip install -e ".[mlip]"
uv pip install --no-deps mace-torchFor Conda environments created from env.yml, run this after activation:
python -m pip install --no-deps mace-torch
python -m pip install -e .MACE-Polar support (--calculator mace-polar) requires the Electrostatic MACE
loader, which is not part of the current PyPI mace-torch release. Install MACE
from the upstream main branch and install graph_electrostatics so the
graph_longrange module is available. IQC defaults this calculator to
model: polar-1-m, device: cpu, and default_dtype: float64; override those
under calculator_params if needed. For the large checkpoint, set
model: polar-1-l, or pass a shared local checkpoint path in
calculator_params.model to avoid repeated downloads in MPI jobs.
On shared HPC systems, limit BLAS/OpenMP threads before importing NumPy,
PyTorch, MACE, or FAIRChem. Otherwise OpenBLAS may try to create one thread
per visible CPU and fail with pthread_create failed:
export OPENBLAS_NUM_THREADS=1
export OMP_NUM_THREADS=1
export MKL_NUM_THREADS=1
export NUMEXPR_NUM_THREADS=1For normal one-process commands, IQC now runs without initializing MPI. This
avoids PMIx/PMI startup failures when running iqc directly on a login or
compute node. For parallel runs, launch IQC through the system MPI launcher:
mpiexec -n 4 iqc --xyz molecules --task singleIf a cluster environment exposes broken PMI/PMIx variables for a serial command, force serial mode:
IQC_DISABLE_MPI=1 iqc --smiles CCC --task opt --calculator uma-s-omolSee docs/calculators.md for calculator names, UMA model mapping, and Hugging Face setup.
The Dockerfile builds a full CPU-only image that bundles IQC with every
open-source quantum-chemistry code and ML potential IQC supports:
xTB, PySCF, NWChem, ExaChem (compiled from source), MACE, FAIRChem UMA, and
ASE EMT. Proprietary codes (ORCA, Gaussian, VASP) are not baked in but can be
mounted at runtime. See docs/docker.md for the full matrix,
proprietary-code mounts, and model-cache setup.
Prefer smaller, purpose-built images? Dockerfile.uv + docker-compose.yml
provide a conda-free, uv-based split family (base / mlip / exachem /
full) so you pull only what a task needs — e.g. docker compose run --rm qc iqc --smiles O --task thermo --calculator xtb. See
Split images with docker compose.
-
Build the Docker image:
docker build -t iqc:full .(The build compiles ExaChem + TAMM from source, so the first build is long; layers are cached for fast rebuilds.)
-
Run the container with Jupyter Lab:
docker run --rm -p 8888:8888 -v "$PWD":/work iqc:fullThen open
http://localhost:8888in your browser. -
Or run the CLI directly (anything after the image name runs in the env):
docker run --rm -v "$PWD":/work iqc:full \ iqc --smiles O --task thermo --calculator xtb -
Verify the image (pytest + a live run of every bundled calculator):
docker run --rm iqc:full bash /opt/iqc/scripts/docker_smoke_test.sh
If you encounter a possible OpenMPI error:
shmem: mmap: an error occurred while determining whether or not /tmp/ompi.yv.1001/jf.0/3074883584/sm_segment.yv.1001.b7470000.0 could be createdTry: export OMPI_MCA_btl_sm_backing_directory=/tmp
PyTorch 2.6 changed torch.load to default to weights_only=True. Older MACE
foundation checkpoints serialize MACE/e3nn/PyTorch model classes, so direct
mace_mp() calls may fail with an error like:
_pickle.UnpicklingError: Weights only load failed
Unsupported global: GLOBAL mace.modules.models.ScaleShiftMACE
IQC's get_calculator(name="mace") applies a compatibility patch that
allowlists the trusted MACE foundation checkpoint classes with
torch.serialization.add_safe_globals(...). If you load MACE checkpoints
outside IQC, use the same safe-global approach for trusted checkpoint sources
or install a MACE/PyTorch combination that handles PyTorch 2.6 safe loading.
IQC provides several computational tasks that can be performed on molecular systems:
single: Single-point energy and force calculationvib: Vibrational frequency calculationir: Infrared spectrum calculationopt: Geometry optimizationthermo: Thermochemical analysisir-thermo: Infrared spectrum and thermochemistry from one shared Hessiannmr: Solution-state NMR shielding and spectrum simulation using ORCA, NWChem, or Gaussian, with optional xTB pre-optimization
These tasks can be specified when running calculations. For example:
from iqc.asetools import run_single_point, run_vibrations, run_optimization, run_thermo
# Single-point calculation
results = run_single_point(atoms)
# Vibrational frequency calculation
results = run_vibrations(atoms)
# Geometry optimization
results = run_optimization(atoms)
# Thermochemical analysis
results = run_thermo(atoms)Each task returns a dictionary containing the results and timing information in milliseconds.
vib, ir, thermo and ir-thermo project translations and rotations out
of the Hessian before extracting frequencies and thermochemistry, for every
calculator (vibration_params.project_trans_rot, default true). The removed
rigid-body contamination is reported in trans_rot_frequencies_cm^-1; see
docs/calculation_recovery.md.
The projector itself is reusable on any ASE Vibrations/VibrationsData
object or raw Hessian via iqc.hessiantools.
Use --calculator to choose the ASE calculator for single, opt, vib,
ir, and thermo tasks:
iqc --xyz molecule.xyz --calculator mace --task opt
iqc --xyz molecule.xyz --calculator uma-s-omol --task single
iqc --xyz molecule.xyz --calculator mace-polar --task singleSupported names are mace, mace-polar, xtb, emt, orca, uma,
uma-s-omol, uma-s-omat, uma-s-odac, uma-m-omol, uma-m-omat, and
uma-m-odac. uma is an alias for uma-s-omol. orca requires the ORCA
executable on PATH (or set ASE_ORCA_COMMAND, or pass command: under
calculator_params in --params).
The ir task computes vibrational frequencies and infrared intensities by
finite-difference displacement. The workflow has three independent steps —
geometry optimization, force evaluation at each displacement (Hessian), and
dipole-moment evaluation at each displacement (intensities) — and each step
can use a different calculator. This is useful when a fast MLIP is good
enough for forces but you want a more accurate electronic-structure method
for dipoles.
Dipole support matters. The dipole calculator must declare 'dipole' in
its implemented_properties. Of the built-in --calculator options, xtb,
orca, and mace-polar do. Regular MACE, EMT, and the UMA models do not, so
they cannot be used as the dipole calculator and cannot drive a
single-calculator IR run.
Use --calculator to apply one calculator to all three steps. Pick a
calculator that supports dipoles:
iqc --task ir --xyz molecule.xyz --calculator xtb--calculator mace (or emt, or any UMA model) will fail at the IR analysis
step with a clear error explaining that the calculator does not implement the
'dipole' property. --calculator mace-polar can be used for IR if your MACE
installation includes the Electrostatic MACE loader.
To compute both IR intensities and thermochemistry without repeating the
optimization and Hessian calculation, use the composite ir-thermo task:
iqc --task ir-thermo --xyz molecule.xyz --calculator xtbir-thermo uses the same IR calculator-role settings described below and then
derives thermochemistry from the vibrational energies produced by that same
finite-difference run.
Per-role calculators are configured through the YAML param file passed with
--params. Any role left unset falls back to --calculator. Calculator names
use the same vocabulary as --calculator (mace, mace-polar, xtb, emt,
orca, uma, uma-s-*, uma-m-*); unknown names are rejected up-front
instead of silently falling back.
A practical mixed workflow uses a fast MLIP for the expensive Hessian and a
DFT calculator for accurate dipoles. ORCA-specific options
(orcasimpleinput, orcablocks, command) go under calculator_params,
which is shared by every role that uses ORCA:
# ir_params.yaml
calculator: mace # fallback for any role not overridden below
calculator_params:
orcasimpleinput: B3LYP def2-SVP
orcablocks: "%pal nprocs 4 end"
command: /path/to/orca # optional; honors $ASE_ORCA_COMMAND otherwise
ir_params:
optimization_calculator: mace
vibration_calculator: mace
dipole_calculator: orcaiqc --task ir --xyz molecule.xyz --params ir_params.yamlThe result dictionary records which calculator was used for each role under
calculator_optimization, calculator_vibration, and calculator_dipole.
The Python API additionally accepts pre-built calculator instances, which is useful when you want different ORCA settings per role (different basis for optimization vs. dipoles), or when wiring in a calculator IQC's CLI does not register:
from ase.calculators.orca import ORCA, OrcaProfile
from iqc.asetools import get_atoms_from_xyz, get_calculator, run_ir
atoms = get_atoms_from_xyz("molecule.xyz")
mace = get_calculator(name="mace")
orca = ORCA(
profile=OrcaProfile(command="/path/to/orca"),
orcasimpleinput="B3LYP def2-SVP",
)
atoms, results = run_ir(
atoms,
optimization_calculator=mace,
vibration_calculator=mace,
dipole_calculator=orca,
)IQC can inspect and process tabular data files with --input / -i.
Inspect a data file without running a calculation:
iqc --input molecules.parquetWhen --input is the only option, IQC prints the file format, row and column counts, schema, null counts, and basic numeric statistics.
Run a calculation from a column containing XYZ text:
iqc --input molecules.parquet --xyz geometry --task singleFor IQC-generated JSONL or parquet files, --xyz can be omitted. IQC then
uses the opt_xyz column by default, which reuses the optimized geometry from
a previous run:
iqc --input iqc_opt_results.parquet --task singleRun a calculation from a column containing SMILES strings:
iqc --input molecules.csv --smiles smiles --task optSort tabular rows before running the task:
iqc --input molecules.parquet --smiles smiles --sort energy --sort_order up --task singleSupported input formats include parquet (.parquet, .pq), CSV/TSV text files, Excel files (.xls, .xlsx, .xlsm, .ods), JSON/JSONL, Feather, and Arrow IPC. For parquet, CSV/TSV, Feather, and Arrow IPC, IQC reads only the requested structure and sort columns where possible.
When --input is used for a calculation, pass one structure selector or omit
both to use the IQC result default:
--xyz COLUMN: read XYZ-format geometry strings fromCOLUMN--smiles COLUMN: read SMILES strings fromCOLUMNand build 3D geometries with RDKit
If neither selector is provided, IQC uses --xyz opt_xyz. This is intended for
rerunning calculations from IQC result files that contain optimized geometries.
Use --sort COLUMN to sort rows before processing. --sort_order up sorts ascending and --sort_order down sorts descending; the default is up. Missing sort values are rejected.
CSV files may contain multiline XYZ values, but those cells must be quoted by the CSV writer. For example:
name,geometry
water,"3
water
O 0 0 0
H 0 0 1
H 1 0 0
"Results written from tabular inputs include input_mode, data_input_file, data_xyz_column or data_smiles_column, data_sort_column, data_sort_order, and data_row_index so each output can be traced back to the source row.
See docs/data_input.md for the full tabular input guide.
When IQC is installed, the Python utilities in scripts/ are available as
console commands:
iqc-jsonl2parquet results.jsonl results.parquet
iqc-read-parquet results.parquet
iqc-optimize-parquet results.parquet optimized.parquet
iqc-reduce-parquet results.parquet --drop 2 5 -o reduced.parquetUse --skip-existing to avoid rerunning calculations that IQC has already
recorded. IQC checks the same calculation identity used by the SQLite unique
key: normalized initial geometry, params, calculator, model, and task.
When --database is provided, existing rows in that SQLite database are used:
iqc --xyz molecules --task single --calculator uma-s-omol \
--database results.db --skip-existingIQC can also index prior IQC JSON/JSONL result files:
iqc --xyz molecules --task single --skip-existing \
--skip-existing-from iqc_single_results_20260501_120000_abcd1234.jsonlIf --skip-existing-from is omitted, IQC scans combined
iqc_*_results_*.jsonl files in the current directory and per-rank JSON files
under tmp/tmp_* result directories. Older top-level tmp_* result directories
are also indexed.
The nmr task runs a backend-aware solution-state NMR workflow for small and medium-sized molecules. It can:
- optimize the input geometry before the shielding calculation
- optionally sample conformers with RDKit and apply Boltzmann weighting
- run GIAO shielding calculations with
orca,nwchem, orgaussian - simulate
1Hand13Cspectra with Lorentzian, Gaussian, or pseudo-Voigt broadening - save plots plus tabulated per-atom and weighted peak data
Reasonable defaults are provided for routine organic-molecule prediction:
- backend:
orca - nuclei:
1Hand13C - geometry optimization: enabled
- conformer sampling: auto-enabled for flexible molecules when RDKit is available
- optimization level:
B3LYP/def2-SVP - shielding level:
PBE0/def2-TZVP - solvent model:
smd - solvent:
chloroform - linewidth:
0.05 ppmfor1H,1.0 ppmfor13C
Built-in reference shieldings are available for 1H and 13C so the workflow works with minimal input, but they are approximate. For production use, pass calibrated values with --reference-shielding.
Minimal ORCA run:
iqc --task nmr --backend orca --xyz molecule.xyz --output-dir nmr_resultsGenerate the starting 3D geometry directly from a SMILES string with RDKit:
iqc --task nmr --backend orca --smiles CCO --output-dir nmr_resultsUse xTB for cheap geometry optimization before ORCA shielding:
iqc --task nmr --backend orca --optimization-backend xtb --xyz molecule.xyz --output-dir nmr_resultsCustomize nuclei, method, basis, solvent, and references:
iqc --task nmr \
--backend gaussian \
--nuclei 1H 13C \
--method B3LYP \
--basis def2-TZVP \
--optimization-method B3LYP \
--optimization-basis def2-SVP \
--solvent-model smd \
--solvent dmso \
--reference-shielding 1H=31.77 \
--reference-shielding 13C=188.10 \
--xyz molecule.xyz \
--output-dir nmr_resultsThe most important CLI flags are:
--backend: shielding backend, one oforca,nwchem,gaussian--optimization-backend: optional optimization backend, one oforca,nwchem,gaussian,xtb--nuclei: nuclei to compute, for example--nuclei 1H 13C--methodand--basis: shielding calculation level--optimization-methodand--optimization-basis: geometry optimization level--solvent-modeland--solvent: implicit solvent settings--chargeand--multiplicity: total charge and spin state--optimize-geometry/--no-optimize-geometry: control geometry optimization--conformer-sampling/--no-conformer-sampling: control RDKit conformer sampling--num-conformers: maximum number of retained conformers--temperature: Boltzmann weighting temperature in Kelvin--linewidth,--lineshape, and--plot-range: spectrum simulation controls--reference-shielding: reference override in the formnucleus=value--output-dir: directory for plots and tables
When --smiles is used, IQC builds a 3D structure with RDKit and starts from the lowest-energy embedded conformer.
xtb is supported only as --optimization-backend. It is not a direct NMR shielding backend.
The workflow writes:
- a combined PNG spectrum plot for the requested nuclei
- one CSV spectrum trace per nucleus
- per-atom CSV and JSON tables with:
- conformer id
- atom index
- element
- nucleus
- isotropic shielding
- converted chemical shift
- conformer energy
- Boltzmann weight
- weighted peak CSV and JSON tables
- a conformer summary CSV with energies, weights, and convergence status
These files make it easy to inspect raw shielding results, compare conformers, and post-process the simulated spectra.