Distil is a system for constructing point-and-click machine learning models, here extended for multi-spectral satellite imagery for timeseries data leveraging an autoML pipeline, adding embedding model trained using self-supervised learning; rapid data labeling facilitated with image query; hierarchical geospatial timeseries modeling; and sub-image feature extraction using weakly supervised segmentation.
2022
2021
-
One of the challenges when building Machine Learning (ML) models using satellite imagery is building sufficiently labeled data sets for training. In the past, this problem has been addressed by adapting computer vision approaches to GIS data with significant recent contributions to the field. But when trying to adapt these models to Sentinel-2 multi-spectral satellite imagery these approaches fall short. To address this deficit, we present Distil, and demonstrate a specific method using our system for training models with all available Sentinel-2 channels.
2018
-
We present in-progress work on Distil, a mixed-initiative system to enable non-experts with subject matter expertise to generate data-driven models using an interactive analytic question first workflow. Our approach incorporates data discovery, enrichment, analytic model recommendation, and automated visualization to understand data and models.
-
This paper describes an abstractive summarization method for tabular data which employs a knowledge base semantic embedding to generate the summary. Assuming the dataset contains descriptive text in headers, columns and/or some augmenting metadata, the system employs the embedding to recommend a subject/type for each text segment. Recommendations are aggregated into a small collection of super types considered to be descriptive of the dataset by exploiting the hierarchy of types in a prespecified ontology.