Full-Time
Domain Scientist – AI Platform
Duration
Full-time
Location
GIFT City, Gandhinagar
careers@iairo.ai
About the Role
We are seeking an experienced Domain Scientist – AI Platform to design, build, and scale biomedical knowledge graphs that power our drug discovery and R&D platforms. In this role, you will work closely with Data Scientists, Software Engineers, AI Engineers, and domain scientists to model complex biological, chemical, and clinical relationships, integrate heterogeneous data sources, and enable graph-driven insights across target identification, drug discovery, and clinical research. You will serve as the bridge between raw biomedical data and structured, queryable knowledge that drives our therapeutic pipelines forward.
Key Responsibilities
- Knowledge Graph Architecture & Development: Design, build, and maintain scalable knowledge graph schemas and ontologies that represent biological entities, chemical compounds, targets, pathways, and clinical data.
- Data Integration & Curation: Ingest, harmonize, and link heterogeneous biomedical and chemical datasets from public and proprietary sources into a unified graph model.
- Cross-Functional Collaboration: Work daily alongside Data Scientists, Machine Learning Engineers, Computational Biologists, and wet-lab scientists to translate drug discovery and clinical research questions into graph-based data models and queries.
- Graph-Powered AI Workflows: Partner with AI Engineers to expose the knowledge graph to agentic and generative AI workflows, enabling automated hypothesis generation, target identification, and evidence retrieval.
- Ontology & Standards Management: Apply and extend biomedical ontologies and standards (e.g., OBO Foundry, UMLS, SNOMED CT) to ensure semantic consistency and interoperability across the platform.
- Query Performance & Scalability: Optimize graph database performance, indexing, and query design to support large-scale, low-latency access for internal and external users.
- Platform Innovation: Contribute to the development of our internal proprietary knowledge graph platform, ensuring it remains robust, reproducible, and accessible to drug discovery teams and external partners.
Essential Expertise and Qualifications
- Education: MSc or PhD in Bioinformatics, Computational Biology, Data Science, or a closely related quantitative discipline.
- Knowledge Graph Experience: Demonstrated hands-on experience building and maintaining knowledge graphs in a biomedical context, such as drug discovery, target identification, or clinical trial research.
- Programming Skills: Strong proficiency in Python and graph query languages (e.g., Cypher, SPARQL, Gremlin). Comfortable writing clean, version-controlled, and reproducible code.
- Graph Technologies: Experience with graph databases and frameworks such as Neo4j, TigerGraph, Amazon Neptune, or RDF triple stores.
- Communication: Exceptional ability to communicate complex data modeling concepts to biologists, and biological/clinical constraints to software engineers and data scientists.
Preferred Qualifications
- Biomedical & Chemical Databases: Familiarity with biological and chemical databases such as Hetionet, PrimeKG, ChEMBL, PubChem, UniProt, DrugBank, Reactome, KEGG, ClinicalTrials.gov, and Open Targets.
- Ontology Frameworks: Hands-on experience with biomedical ontologies and standards, including Gene Ontology (GO), Human Phenotype Ontology (HPO), MeSH, UMLS, and SNOMED CT.
- AI/ML Integration: Experience integrating knowledge graphs with machine learning pipelines, embeddings (e.g., knowledge graph embeddings, GNNs), or large language model–based retrieval systems.
- Clinical & Drug Discovery Context: Prior experience supporting target identification, drug repurposing, mechanism-of-action analysis, or clinical trial research using graph-based approaches.
- Cloud & Infrastructure: Experience deploying graph pipelines in cloud environments (AWS, GCP, Azure) and working with containerized workflows (Docker, Kubernetes, Nextflow, or Snakemake).
- Data Engineering: Familiarity with ETL/ELT pipelines and semantic data integration tools (e.g., RDF/OWL tooling, Apache Airflow, or similar orchestration frameworks).
What We Offer
- Opportunity to work at the absolute cutting edge of AI and biomedical data, driving real-world impact in therapeutic discovery.
- A collaborative, talent-dense team of world-class scientists and engineers.
- Competitive compensation package, including equity and comprehensive health benefits.
- Highly productive work environment with state-of-the-art computational resources including GPU clusters.
JOIN OUR TALENT NETWORK
Don't see the right role?
We're always looking for exceptional researchers and engineers. Reach out directly.