-
Notifications
You must be signed in to change notification settings - Fork 4
Expand file tree
/
Copy path11-pathology-python.Rmd
More file actions
320 lines (225 loc) · 12.7 KB
/
Copy path11-pathology-python.Rmd
File metadata and controls
320 lines (225 loc) · 12.7 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
# Working with digital pathology data and python
Histopathology is used to analyse tissues from the body under the microscope. When a sample of tissue is taken during and operation or biopsy, it is put into a preservative, dehydrated and then impregnated with paraffin wax. It may also be frozen. This allows the tissue to be very thinly sliced and mounted onto glass slides, where the tissue can then be stained and studied under the microscope. Different types of stains can be used to look at different things.
These stained slides can be scanned with a slide scanner to allow the image to be digitised using scanners. There are two main types of images. Brightfield - generated by using white light, which requires deconvolution to produce separate colour channels. Fluorescence - generated using a light source to excite fluorescent probes, this typically generates images with 2+ colour channels.
The file formats produced by these scanners varies by manufacturer and scanning technique. Essentially, data is stored as tiled TIFF images within what is possibly an XML structure. There are usually multiple images produced, not only of the tissue but also of the slide, including a macro image (no magnification, sometimes helps scanner locate tissue), and the slide label.
Unfortunately, it's quite a difficult format to interrogate easily. However, there is a nice program written that can solve this issue, known as [OpenSlide](https://openslide.org/).
OpenSlide has C, Python and Java versions. The python version is probably easiest to interact with, either directly in Python, or it can be combined with R using `reticulate`.
This chapter will not go into depth about issues around environments or how to install python properly. Please familiarise yourself with how to use `pip`, particularly `pip install` and set / activate python environments.
You can use RStudio for working in python too and this is what is suggested here. Other formats exist including jupyter notebooks.
For more information, [Python for microscopists](https://github.com/bnsreenu/python_for_microscopists) has excellent resources.
## Installing OpenSlide
First install the linux version of openslide - the python files act as a wrapper to this system library.
```{sh, eval = FALSE}
apt-get install python3-openslide
```
To install openslide and the libraries we will need, first create a python environment. Do this inside the folder you are going to be working in.
```{sh, eval = FALSE}
virtualenv -p python3 openslide_env
```
You can then either activate the virtual environment to install python libraries there, or install them system/user wide.
```{sh, eval = FALSE}
python3 -m pip install openslide-python
```
We will also need some other libraries
```{sh, eval = FALSE}
python3 -m pip install numpy matplotlib tifffile
```
Installation is fairly straighforward -if there are issues, do it as a sudo user. Most of the problems will appear if you are trying to use different or conflicting versions of python or if there are incorrect permissions for installation.
## Using OpenSlide
Once openslide is installed, we can now write python scripts. Do this in RStudio or other IDE in the same folder as your virtual environment. You can create a project in RStudio using this existing folder, then add new python scripts. Python scripts have the suffix `.py`.
First, start by checking you can successfully import the libraries.
```{python, eval = FALSE}
from openslide import open_slide
import openslide
from PIL import Image
import numpy as np
from matplotlib import pyplot as plt
import tifffile as tiff
```
If you get an error- try installing the libraries with `pip` as above.
### Opening a slide and inspecting properties
```{python, eval = FALSE}
file_path = '/home/test/slides/myslide.svs' # replace with where your slides are
slide_in = open_slide(file_path)
slide_in_props = slide_in.properties
slide_in_assc_images = slide_in.associated_images.items() # see if there are any associated images
```
## Whole workflow
Here is the whole workflow for extracting magnified tiles of the image, taken from [Python for microscopists](https://github.com/bnsreenu/python_for_microscopists).
```{python, eval = FALSE}
file_path = '/home/test/slide_ins/myslide_in.svs' # replace with where your slide_ins are
slide_in = open_slide(file_path)
slide_in_props = slide_in.properties
print(slide_in_props)
print("Vendor is:", slide_in_props['openslide.vendor'])
print("Pixel size of X in um is:", slide_in_props['openslide.mpp-x'])
print("Pixel size of Y in um is:", slide_in_props['openslide.mpp-y'])
#Objective used to capture the image
objective = float(slide_in.properties[openslide.PROPERTY_NAME_OBJECTIVE_POWER])
print("The objective power is: ", objective)
# get slide_in dimensions for the level 0 - max resolution level
slide_in_dims = slide_in.dimensions
print(slide_in_dims)
#Get a thumbnail of the image and visualize
slide_in_thumb_600 = slide_in.get_thumbnail(size=(600, 600))
slide_in_thumb_600.show()
#Convert thumbnail to numpy array
slide_in_thumb_600_np = np.array(slide_in_thumb_600)
plt.figure(figsize=(8,8))
plt.imshow(slide_in_thumb_600_np)
#Get slide_in dims at each level. Remember that whole slide_in images store information
#as pyramid at various levels
dims = slide_in.level_dimensions
num_levels = len(dims)
print("Number of levels in this image are:", num_levels)
print("Dimensions of various levels in this image are:", dims)
#By how much are levels downsampled from the original image?
factors = slide_in.level_downsamples
print("Each level is downsampled by an amount of: ", factors)
#Copy an image from a level
level3_dim = dims[2]
#Give pixel coordinates (top left pixel in the original large image)
#Also give the level number (for level 3 we are providing a valueof 2)
#Size of your output image
#Remember that the output would be a RGBA image (Not, RGB)
level3_img = slide_in.read_region((0,0), 2, level3_dim) #Pillow object, mode=RGBA
#Convert the image to RGB
level3_img_RGB = level3_img.convert('RGB')
level3_img_RGB.show()
#Convert the image into numpy array for processing
level3_img_np = np.array(level3_img_RGB)
plt.imshow(level3_img_np)
#Return the best level for displaying the given downsample.
SCALE_FACTOR = 32
best_level = slide_in.get_best_level_for_downsample(SCALE_FACTOR)
#Here it returns the best level to be 2 (third level)
#If you change the scale factor to 2, it will suggest the best level to be 0 (our 1st level)
#################################
#Generating tiles for deep learning training or other processing purposes
#We can use read_region function and slide_in over the large image to extract tiles
#but an easier approach would be to use DeepZoom based generator.
# https://openslide.org/api/python/
from openslide.deepzoom import DeepZoomGenerator
#Generate object for tiles using the DeepZoomGenerator
tiles = DeepZoomGenerator(slide_in, tile_size=256, overlap=0, limit_bounds=False)
#Here, we have divided our svs into tiles of size 256 with no overlap.
#The tiles object also contains data at many levels.
#To check the number of levels
print("The number of levels in the tiles object are: ", tiles.level_count)
print("The dimensions of data in each level are: ", tiles.level_dimensions)
#Total number of tiles in the tiles object
print("Total number of tiles = : ", tiles.tile_count)
#How many tiles at a specific level?
level_num = 11
print("Tiles shape at level ", level_num, " is: ", tiles.level_tiles[level_num])
print("This means there are ", tiles.level_tiles[level_num][0]*tiles.level_tiles[level_num][1], " total tiles in this level")
#Dimensions of the tile (tile size) for a specific tile from a specific layer
tile_dims = tiles.get_tile_dimensions(11, (3,4)) #Provide deep zoom level and address (column, row)
#Tile count at the highest resolution level (level 16 in our tiles)
tile_count_in_large_image = tiles.level_tiles[16] #126 x 151 (32001/256 = 126 with no overlap pixels)
#Check tile size for some random tile
tile_dims = tiles.get_tile_dimensions(16, (120,140))
#Last tiles may not have full 256x256 dimensions as our large image is not exactly divisible by 256
tile_dims = tiles.get_tile_dimensions(16, (125,150))
single_tile = tiles.get_tile(16, (62, 70)) #Provide deep zoom level and address (column, row)
single_tile_RGB = single_tile.convert('RGB')
single_tile_RGB.show()
###### Saving each tile to local directory
cols, rows = tiles.level_tiles[16]
import os
tile_dir = "images/saved_tiles/original_tiles/"
for row in range(rows):
for col in range(cols):
tile_name = os.path.join(tile_dir, '%d_%d' % (col, row))
print("Now saving tile with title: ", tile_name)
temp_tile = tiles.get_tile(16, (col, row))
temp_tile_RGB = temp_tile.convert('RGB')
temp_tile_np = np.array(temp_tile_RGB)
plt.imsave(tile_name + ".png", temp_tile_np)
```
## Using python and R together
Using the `reticulate` package, R can communicate with python. This is slightly glitchy at the time of writing. So writing python objects from R seems to not work (using `py$` as mentioned in `reticulate` documentation), but python calling objects in the R environment seems reliable.
To do this we simply access the R environment in python using `r['myobject']`. Lets use the `file_path` example, where you might have used R to save the file path of specific slide file, then you want to use python to extract or manipulate the images.
In R:
```{r, eval = FALSE}
file_path = '/home/test/slides/myslide.svs' # replace with where your slides are
```
In python:
```{python, eval = FALSE}
file_path = r['file_path']
slide_in = open_slide(file_path)
```
## Handling errors in python
Most slides will be scanned well, newer scanners use lasers and a variety of different tools to automatically detect where tissue is and the depth at which to scan. However, some files will not scan properly or be corrupted in transfer (files are at least 200M+ per scan, ususally around 500M).
If there are errors, it is quite annoying, particularly given the size of files being handled if a run fails due to errors in opening. To avoid this we can use the python `try` commands, where it will attempt the command but continue on fail.
```{python, eval = FALSE}
#import pyvips
from openslide import open_slide
import openslide
from PIL import Image
import numpy as np
from matplotlib import pyplot as plt
import tifffile as tiff
#import
file_path = r['file_path']
try:
slide_in = open_slide(file_path)
slide_in_props = slide_in.properties
r['slide_properties'] = slide_in_props
r['slide_associated_images'] = slide_in.associated_images.items()
except Exception as e:
r['slide_properties'] = print("Error loading") # doesn't save error
r['slide_associated_images'] = print("Error loading") # doesn't save error
```
## Iterating over many many slide files
Combining all the above, we can massively extend functionality and the number of slides we can process, for example to extract hundreds of slides at once and apply the same analysis to them.
In R:
```{r, eval = FALSE}
library(tidyverse)
library(reticulate)
Sys.setenv(RETICULATE_PYTHON = "/home/slides/openslide_env/bin/python")
#list files in a dir
files = data.frame(file_path = list.files('/home/slides',
pattern = '.svs|.scn',
full.names = T, recursive = T))#scn needs different
#make names
files = files %>%
mutate(file_name = basename(file_path),
label_path = gsub('.svs|.scn', '', file_name),
out_path = gsub('.svs|.scn', '_label.png', file_path),
file_info = file.info(file_path))
#first get label properties
slide_properties_list = list()
associate_images_list = list()
#get properties
for(i in 1:length(files$label_path)){
label_path = files$label_path[i]
file_path = files$file_path[i]
out_path = files$out_path[i]
source_python('label_properties.py', envir = globalenv())
slide_properties_list[[i]] = as.character(slide_properties)
associate_images_list[[i]] = as.character(slide_associated_images)
}
openslide_data = data.frame(do.call(cbind, list(files$file_path, associate_images_list, slide_properties_list)))
```
In python `label_properties.py`:
```{python, eval = FALSE}
#import pyvips
from openslide import open_slide
import openslide
from PIL import Image
import numpy as np
from matplotlib import pyplot as plt
import tifffile as tiff
#import
file_path = r['file_path']
try:
#slide_in = open_slide('/mnt/data/tdrake/light_microscopy/immune_staining_from_colin/01082019/collection_0000020061_2019-08-01 13_05_34.scn')
slide_in = open_slide(file_path)
slide_in_props = slide_in.properties
r['slide_properties'] = slide_in_props
r['slide_associated_images'] = slide_in.associated_images.items()
except Exception as e:
r['slide_properties'] = print("Error loading")
r['slide_associated_images'] = print("Error loading")
```