Company
Portfolio Data
INSILICA LLC
UEI: RTN8V2BMGY63
Number of Employees: 2
HUBZone Owned: Yes
Woman Owned: No
Socially and Economically Disadvantaged: No
SBIR/STTR Involvement
Year of first award: 2020
5
Phase I Awards
1
Phase II Awards
20%
Conversion Rate
$1,321,085
Phase I Dollars
$1,014,776
Phase II Dollars
$2,335,861
Total Awarded
Awards
srvc-EHSR: Sysrev Version Control - Environmental and Health Related Systematic Reviews
Amount: $1,014,776 Topic: R
While document reviews (DRs) are not the only requirement of environmental and health-related risk assessments (EHRRAs), this component contributes considerable time and money as each DR costs >$140,000. In the pharma industry alone, average yearly DR expenditure is $16.7M. Currently, most EHRRAs rely on manual curation of data including selection of target documents & parsing key data fields in sources. This process, while integral to accuracy & completeness of EHRRAs, requires hundreds of research hours. While increasing numbers of platforms utilize, or permit utility of, machine learning (ML) models, their complex codebase makes it too cumbersome to deploy new features or integrate ML tools for specific use cases. As a result, risk assessments continue to rely on long, manual document reviews which, in turn, delays accessibility to new data for informed chemical & product safety decisions. Therefore, academia, industry, & government regulators would benefit from a platform leveraging ML to optimize EHRRA related DRs. Insilica, LLC will solve these market challenges via a cloud-based or on-site hosted, application, which is remote-capable and enables EHRRA researchers to build customizable workflows based on standardized, interchangeable DR & ML components. Using the intuitive srvc-EHSR interface, researchers will use natural language inputs to define inclusion/exclusion parameters for an environmental chemical review and required data summary statistics. A list of publications or documents will be returned, along with summarized data metrics for defined fields and research questions across all sources. A flexible algorithm structure will enable Insilica’s proprietary ML tools to be leveraged for reviews, as well as customizable open-source tools be shared across researchers and organizations. Once a custom defined DR is complete, individual sub-components of the review can easily be repurposed for future DRs or shared with new researchers to accelerate longitudinal open-source knowledge aggregation in targeted chemical and environmental human health hazard applications. This project will build on a successful Phase I, in which 1) modular data architecture defined, captured, & synchronized data across varying form factors of literature & documents, 2) front end query coupled with ML tools enabled automatic selection of document sources & data summaries with accuracy similar to expert researchers, and 3) a usability study reported high researcher satisfaction. This Phase II will expand the foundation to a commercially viable product with fully validated software, ML tools, & support infrastructure. First, architecture will be enhanced by cross platform interoperability, mobile responsiveness, and self-hosting. Next, ML accuracy will be improved, and open-source options integrated. User front end tools will also be expanded with onboarding, collaboration, administrator, & reporting options. After components are integrated at the system level, migrated to production server, and tested to compliance standards, it will be deployed in a field study to document technical performance, efficiency, & cost savings.
Tagged as:
SBIR
Phase II
2025
HHS
NIH
srvc-EHSR: Sysrev Version Control - Environmental and Health Related Systematic Reviews
Amount: $293,741 Topic: R
Environmental & Health-Related Risk Assessments (EHRRA) are an integral part of new product formulation, as well as product safety stewardship. Unfortunately, the current process for conducting EHRRAs is expensive and slow. According to the EPA report: "FY 2021 Contributions to EPA's Portfolio of Evidence", the cost of conducting a Toxic Substance Control Act (TSCA) Chemical Evaluation is approximately $8.4MM over 3.5 years. Currently, most EHRRAs rely on the manual curation of data via Systematic Literature Review. This process, while integral to the overall accuracy and completeness of EHRRAs, requires hundreds of research hours. Therefore, academia, industry, and government regulators would all likewise benefit from a platform which leverages machine learning to optimize EHRRA related literature and systematic reviews. More specifically, there is currently a strong value proposition in the $534BB global cosmetic industry, $221BB global agrochemical industry, $235BB global home cleaning supplies industry, and $1.5TT global consumer packaged goods industry for tools that optimize SLRs for EHRRAs. To minimize the amount of time required by SLRs for EHRRAs, advances in machine learning (ML) are increasingly helping researchers more quickly identify relevant pieces of information. One major obstacle to the development and use of additional ML tools for EHRRAs is integrating such tools into existing SLR platforms. While an increasing number of these platforms utilize, or permit the utility of, ML models, their complex codebase makes it too cumbersome to deploy new features or integrate ML tools for specific use cases. As a result, risk assessments continue to rely on long, manual literature reviews which, in turn, delays accessibility to new data to make informed chemical and product safety decisions. The 'srvc-EHSR' platform will solve this growing market need though a git-based, remote-deployment capable, application which enables EHRRA researchers to build customizable workflows based on standardized, interchangeable SLR and ML components. The platform will prioritize internal and external interoperability of components to ensure fit-for purpose adaptability. As a result, srvc-EHSR will enable more efficient SLRs and more thorough EHRAAs, thereby decreasing the time-to-market for new products and increasing the overall safety of consumer and industrial chemicals. While the fully commercialized srvc-EHSR will integrate the features above, Phase I will target feasibility for component modularity and interchangeability, ML development, and prototype interface. Development will leverage existing assets, BioBricks.ai and Sysrev, to develop core Phase I srvc-EHSR components as cost-efficiently as possible. A prototype cloud based application, terminal interface, software “packages”, and API will be developed and deployed in a usability study by the Phase I Commercial Partner wherein researchers will use srvc-EHSR to review documents for environmental, chemical, and health data.
Tagged as:
SBIR
Phase I
2024
HHS
NIH
BioBricks-Env: AI driven, open-source, modular, composable platform to organize, store, retrieve, extract, and integrate environmental health & risk related data
Amount: $295,662 Topic: R
The overall global market for toxicology testing was $8.1 billion in 2019 and is expected to reach $27 billion by 2025. Risk assessment, which historically has relied on toxicology testing, is an integral process of new product development and ongoing product safety stewardship. When answering bioinformatics questions, researchers often struggle to discover appropriate sources from hundreds of public resources, especially when each resource provides its own custom distribution method. To minimize time to market and expense, advances in computational biology and machine learning (ML) are helping to automate and reduce the expense of chemical testing via New Approach Methodologies. Scientists are leveraging public knowledgebases, which catalog chemical, biological, and toxicity data, to increase the scope & performance of advanced computational tools. There is currently a strong value proposition in the $511B cosmetic industry, $218B agrochemical industry, $221B home cleaning supplies industry, and $2T Consumer Goods markets for tools that ensure environmental safety & expedite product design strategies by allowing researchers to rapidly leverage all available information. BioBricks-Env will solve this growing market challenge by providing a standard protocol supported by “bricks” and an open-source tool set. Bricks are stand-alone digital resources that can be easily loaded into any data- science environment. BioBricks-Env will leverage git & data-version-control to create “data-dependencies”, data resources that can be imported into a data science environment using methods similar to package management tools like Homebrew. However, BioBricks-Env will relieve developers of the significant effort and burden of maintaining data pipelines, and if adopted, relieves data developers of the cost of maintaining a distribution method. BioBricks thereby provides a service to data scientists, who can reduce the cost of project development, and data distributors, who may struggle to find adoption due to the lack of an easy distribution method. This service fills a strongly needed niche in the public health data marketplace. Revenue will be driven through two sources including 1) Software as a Service models to toxicology researchers and laboratories and 2) strategic partners who monetize their proprietary data on the BioBricks-Env platform. While the fully commercialized BioBricks-Env will integrate all features above, Phase I will target feasibility for standardizable brick architecture, simplicity of deployment and use, interoperability of outputs, AI for data sourcing, and prototype interface. Development will focus on public databases of high relevance to environmental health and risk data. The Phase I BioBricks-Env prototype will provide a data-dependency package manager that can run on windows, mac, and Linux and support common “Install”, “Update”, and “Load” concepts. AI tools to support data integration and mapping will be expanded to support environmental data. Once the prototype has been developed and tested in house, it will be evaluated in a usability and cost- performance validation study with a commercial beta testing partner.
Tagged as:
SBIR
Phase I
2024
HHS
NIH
ToxIndex-CPG: Machine learning driven platform integrating a hazard susceptibility database to quantify chemical toxicity factors, predict risk levels and classify biological responses
Amount: $255,880 Topic: R
The overall global market for toxicology testing was $8.1 billion in 2019 and is expected to reach $27 billion by 2025. As toxicological testing is a pre-requisite step in most product development, it adds significant time and costs, as well as represents human health hazards when key data is not captured. In order to minimize time to market, expense, and animal use, advances in computational biology and machine learning (ML) are helping conduct more efficient in-silico simulations. These strategies are driving strong growth for advanced computational tools. More specifically, there is currently a strong value proposition in the $635 billon CPG market for tools that ensure safety and expedite product design strategies by linking toxicology hazard profiles in reproductive health to chemicals, exposure and product use cases. This will allow a better understanding and prioritization of chemicals for integration in products to minimize associated reproductive health hazards.The ToxIndex-CPG platform will solve this growing market need through a web-based interface that allows CPG toxicology researchers access to customized data for early product planning and study design. The platform will focus on continuous curation of a database to maintain known relationships in existing literature and data sources, as well as advanced algorithms for predictive relationships for unknown combinations. This project will target CPG products and reproductive health hazards, as this represents major markets and risks to vulnerable populations. The user front end will be designed as a web-based tool for toxicology researchers to query specific chemicals, CPG use cases, and health hazards. Based on query inputs, the platform will return a sorted and ranked list of potential adverse reproductive health outcomes. Researchers will be able to explore impact of specific chemicals on ranked reproductive hazards through advanced visualization tools. Hazard relationships between chemicals and human factors and planned CPG product use cases will be learned through ML using quantitative structure-activity relationship (QSAR) models. The platform will leverage existing data sources for chemical and medical data to build models and continue to adaptively learn as datasets continue to grow. The platform will prioritize application programming interfaces (API) to support a growing market of cheminformatics developers.Phase I will target feasibility of data aggregation, ML development, and prototype interface. Development will leverage an existing tool, Sysrev, for automated data extraction from publications and data sources to increase likelihood of success. The Sysrev platform will parse existing data sources to extract known human factors and use case susceptibility factors for a given chemical toxicant and reproductive health hazards. This will create an initial hazard database of known factors as a gold standard for ML testing. Next, QSAR ML models will be developed to associate chemicals to hazards, and then mediation models from chemicals through hazards to understand causality likelihood in specific human factors and use cases of those chemicals. Finally, a prototype web app and visualizations will be developed and deployed in a usability study with toxicology market users.
Tagged as:
SBIR
Phase I
2022
HHS
NIH
Survive-OUD: AI Platform to Integrate Complex Data Sources, Predict Relapse, and Recommend Interventions for Opioid Use Disorder
Amount: $251,348 Topic: NIDA
The current opioid crisis is significantly impacting millions of lives, healthcare, social welfare, and the economy. Patient interactions with the treatment system coupled with results of completed studies create a wealth of data stored in disparate electronic sources. Therapists and health care providers with limited time and resources face challenges to access, integrate and monitor this vast data set for novel opportunities to improve care. Significant advances would include predicting when a patient will relapse out of a program, and also suggesting optimal personalized care strategies to reengage patients before this negative event occurs.Survive-OUD will solve these challenges by providing a web-based therapist interface integrating survivor model artificial intelligence (AI) strategies. Based on multiple input data domains leveraged from existing electronic medical record sources, survivor recurrent neural networks will be trained to recognize when patients are likely to relapse or drop out of an OUD program. Examples of data domains that can be input to the network include patient demographics, medical and prescription data, engagement with therapy paradigms, and compliance with logistical program tasks. Furthermore, once a patient is noted as high risk, a second layer of algorithms will be developed to recommend a specific and personalized care strategy for retention based on existing best practices in the literature and clinical trials. Therefore, the Survive-OUD platform will also integrate with common literature database and clinical trial repositories. Utilizing an existing AI platform for searching, tagging, and extracting data from database sources, the innovative platform will close the loop on actionable results by recommending updated care options based on potential outcomes learned from best practices in existing literature. The AI architecture developed will greatly improve success rates in opioid addition programs and expand high quality healthcare.While the commercialized Survive-OUD platform will integrate all features above, Phase I will target feasibility of data aggregation and AI algorithms to detect relapse and recommend intervention strategies. The innovative technical challenge in Phase I is to develop and validate targeted AI tools using data already being captured in patient workflow to allow early prediction of patient retention issues. More specifically, a prototype therapist interface and data network infrastructure will be developed to source personalized patient data as well as literature and clinical trial sources. Once the platform architecture has passed verification testing, it will be deployed in a field data collection study to determine usability and also provide a rich set of de-identified data for algorithm development. Collected data will then be used to train and test AI algorithms for early detection of patient dropout/relapse and appropriate treatment recommendation.The objective is to design, develop, and demonstrate feasibility of Survive-OUD, an artificial intelligence driven platform to integrate complex data sources, predict patient relapse, and recommend intervention strategies for individuals impacted by opioid use disorder. Currently millions of Americans suffer from an opioid use disorder (OUD) and program relapse rates are extremely high. Therefore, an advanced, bioinformatics platform that accurately predicts when OUD patients will drop out of programs and offers personalized prevention strategies would provide a novel clinical tool to help combat the opioid crisis in the U.S.
Tagged as:
SBIR
Phase I
2020
HHS
NIH
SBIR Phase I: Advanced Cancer Analytics Platform for Highly Accurate and Scalable Survival Models to Personalize Oncology Strategies
Amount: $224,454 Topic: DH
The broader impact/commercial potential of this Small Business Innovation Research (SBIR) Phase I project will develop personalized clinical decision-making in cancer care. An estimated 17 million cases of cancer are diagnosed globally each year. Over $90 billion per year is spent in total on cancer-related health care in the U.S., and cancer patients pay over $4 billion out of pocket for health care. Therapeutic strategy selection and clinical trial research targeted to oncology become exponentially complex when unique types of cancer are considered, as well as how they may uniquely impact gender, race, ethnicity, and age of affected populations. The proposed technology will develop advanced bioinformatics models and visualization tools to guide decision-making by oncologists. It will develop and use advanced survival models targeting cancer types, other biological and chemical factors, and patient demographics. This Small Business Innovation Research (SBIR) Phase I project will focus on three objectives. 1) We will develop and validate transfer learning models that leverage large data sets from high-incidence cancer types to improve results of cancer types with sparse data. 2) We will leverage these data in a disease-agnostic platform using a recurrent neural network to account for temporal variation to predict survivability. 3) We will develop visualization tools for clinicians to understand causal relationships. This system will use several innovations: a) Transfer Learning to Scale Available Data: Since cancer survival modeling is limited in many cancer types due to lack of data, we will demonstrate the feasibility of transfer learning in this context. b) Single Recurrent Neural Network: We will implement a recurrent neural network to improve performance and allow a single network to be trained across all cancer types and patient population characteristics. c) Control Feature Mediation Analysis: We will develop accurate survival models with an understanding of the sensitivity to inputs. d) Clinician-Driven Interpretation and Visualization Tools: The framework needs interpretation and visualization features to reduce data into reports easily digestible for clinical decision-making. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
Tagged as:
SBIR
Phase I
2020
NSF