Wednesday, 22 July 2015

OSM Retail Survey: Part-4

It is often useful to raise our sights from the data that has been recorded in OSM, and consider the data that hasn't been recorded. As discussed above, the overall level of retail coverage in England is 27% of retail premises. However, there are variations in the extent to which different types of shop are recorded.

There are a number of ways to assess OSM coverage of a specific sector at a national level. The approach, broadly, is to count the number of shops of a particular type that are already recorded in OSM, estimate the total number that should be there, and compare the two.

There are various sources of statistics that give basic information on numbers of different types of shop at a national level. Getting useful estimates of numbers at a more local level is a bigger challenge.


  • It isn't difficult to find information on the number of branches for major retail chains - through their own publications, from business reporting, or from Wikipedia. The OSM Wiki has a detailed page on major UK Retail chains. 
  • There are national statistics which can be used for some sectors and types of shop (for example, see UKBA01a Enterprise/local units by 4 Digit SIC and UK Regions; and Retail Hereditaments by Administrative Area issued by the Valuation Office Agency). 
  • For independent specialists, many trade associations publish figures on the size of their sector.
  • Press articles and market research companies will sometimes publish figures on the number of different types of retailer. 
  • When all else fails, searching a directory  (such as Yellow Pages) can give some idea of the likely number of outlets. 
  • Where retailers need to be licensed (tattoo parlours, for example), I thought it would be easy to obtain figures at a local authority level. No doubt this would be possible (through FOIA, for example), but so far I haven't found a more accessible source of licensing statistics. Local figures may be easier to find from the licensing authority, and any ideas would be welcome on where to find national figure.

For specific retail locations:


  • within the OSM community, Robert Whittaker has tools relating to Post Offices on his Post Hoc pages (http://robert.mathmos.net/osm/postboxes/). 
  • Beyond the community, most large retail chains publish a list of branches, and some have given permission for this information to be added to OSM. 
  • Trade associations for specialist independent retailers don't normally seem to provide information on the location of individual members, but some may,
  • The NHS provides lists of pharmacies and opticians.

Using this approach, and taking three examples where we might expect levels of recording to be relatively high:

  • There are 11,696 post offices in the UK, and I have found 7,622 (65%) of them in OSM. Around 90% of built-up areas with a population of more than 5,000 have a post office in OSM. The 72 that don't might be good places to find missing post offices. More generally, post offices can be used as an indicator of a wider gap in coverage. They are one of the types of retail outlet that are likely to be added before other retail properties. So a larger settlement with a missing post office is likely to contain other retail properties that need to be added. 
  • There are 11,647 community pharmacies in England. I found 4,225 tagged as a pharmacy (36% of the total), and another 483 tagged as “chemist”. Strictly speaking “chemist” is for shops that don't supply prescriptions, but has been quite widely used as a synonym for “pharmacy”. Taken together we locate about 40% of community pharmacies. More than half are missing. I imagine that towns the size of Braintree, Grantham, Peterlee, Melton Mowbray, Haverhill, Maghull, and Congleton have a community pharmacy – but none seems to be recorded in OSM (they are displayed in Google location searches). Only around half of the pharmacies in England are specialised shops – the rest are an operation embedded within another store. Community pharmacies that operate from within a large supermarket seem to be under-recorded in OSM. 
  • There are about 2,500 specialist bicycle shops in the UK, and I have found 1,631 (65%) in the OSM database. The largest bicycle retailer, Halfords, has 465 branches across the UK, of which I found 364 (78%). I'm not sure how many of those offer bicycles, but OSM says that 151 of them do (41%). That must surely under-state the true figure.

And some examples where I expected coverage to be relatively low:

  • Figures suggest that there are about 1,400 pound stores in the UK. I've found 586 tagged with “variety_store”, and another 170 or so with alternative tagging. Which means that tagging is inconsistent, but suggests that coverage is over 50% - i.e. more than I expected. Perhaps my estimate of the total is too low
  • There are 8,500 Charity Shops in England, 900 in Scotland and 500 in Wales. I should find 9,900 in my data extract. Depending on how carefully I interpret the data, I can find between 1,751 and 1,994 (18-20%). Around 90% are tagged as “shop=charity” but there is a smattering of others tagged according to their specialisation: “shop=clothes”, “shop=secondhand” or “shop=books”



Primary
tag value
OSM count
(UK)
OSM count
(England)
Estimated 
 actual (UK)
Estimated 
 actual (England)
Approx. coverage
pub
34,937
31,180
48,000

73%
restaurant
16,062
13,855
60,000

27%
fast_food
15,762
13,794

41,295
33%
cafe
13,137
11,280
16,501

80%
convenience
13,108
11,212
48,303

27%
supermarket
8,720
7,352
6,410

119%
post_office
7,622
6,199
11,696

65%
hairdresser
7,187
6,366
38,300

19%
fuel
6,207
5,190
8,588

72%
bank
5,946
5,089
8,961

66%
pharmacy
4,871
4,225

11,647
36%
charity
1,682
1,476
9,900
8,500
17%
bicycle
1,631
1,406
2,500

65%
beauty
1,543
1,361
13,000

12%
bookmaker
1,386
1,242
9,128

15%
optician
1,161
1,041
7,250

16%
florist
983
885
8,000

12%
alcohol
784
665
5,575
4,195
14%
variety_store
586
528
1,400

42%
deli
456
391
2,500

18%
seafood
87
70
950

9%



It is interesting to consider in more detail at how data users might interpret some specific examples.

Finding a pharmacist (i.e. someone who can dispense prescriptions) could be the basis of a useful application, and there have been various attempts to develop appropriate tagging, but the actual data is quite complex for data users to interpret.

Values of “pharmacy” and “chemist” can appear for “amenity” and “shop”; “dispensing” can be set to “yes”, “no” or sometimes the name of the outlet. And all of these can be combined in different ways, alongside other values of “amenity” and “shop”.

  • “amenity=pharmacy” alongside any value of “shop=*” and either “dispensing=yes” or no value for “dispensing”: this is in line with the various guidelines, and unambiguously indicates that prescriptions will be dispensed. This accounts for almost 90% of cases in the data
  • “shop=chemist” without any indication of “dispensing”: is correct tagging for a place where prescriptions will NOT be dispensed, but examining actual examples suggests that it is widely mis-used for pharmacies. So in practice it has to be regarded as ambiguous. It represents almost 10% of cases.
  • “amenity=pharmacy” with “dispensing=no”: is inconsistent tagging, and not in line with the guidelines, but can still be interpreted fairly confidently as a place where prescriptions will NOT be dispensed. It accounts for around 1% of cases
  • “shop=chemist” without “amenity=pharmacy”, and with “dispensing=no” is correct tagging, and unambiguously a place where prescriptions will NOT be dispensed. It only accounts for 0.1% of cases.
  • “shop=pharmacy”, “amenity=chemist”, with or without other values are examples of incorrect use of the tags, but small in volume (less than 0.5%), and often appear alongside a correct tag (e.g. “shop=pharmacy”+“amenity=pharmacy”): the incorrect tag values can safely be ignored by data users without sacrificing significant amounts of relevant data

The above figures are calculated from pharmacies and chemists recorded in the database. So it is worth recalling that this only accounts for 40% of actual pharmacies, and around 60% of these outlets do not appear in the database at all.

In practice data users are going to be reasonably confident that they have found a dispensing pharmacist where “amenity=pharmacy” is present, and “dispensing” is either absent, or set to anything other than “no”. They will have to treat “shop=chemist” as ambiguous in this context. In practice they will probably ignore everything else because the complexity of the logic increases out of all proportion to the quantity of reliable data that it can uncover. In summary they will confidently interpret 90% of the data in the database, and find just over one in three pharmacies. If they interpret the data more loosely they will be able to point their users to about 40% of real pharmacies. If they want to find more, then at present they will have to look elsewhere for their data.

Next we will look at how coverage by type of retail outlet might be used to provide useful feedback to contributors and data users.....

OSM Retail Survey: Part-3

It is quite difficult to identify areas of incomplete retail mapping within a large conurbation.

Initially I thought investigation of areas marked as “landuse=retail” showed promise, but they proved disappointing in practice. The meaning of this tag has been interpreted in different ways, and used inconsistently. Almost half of retail areas in OSM contain no retail properties, so they can provide some indication of gaps. However, some of them are very small, and two-thirds of retail properties already in OSM lie outside a marked retail area, so in general checking for data on shops within retail areas is unlikely to be efficient.

  • A number of fairly large settlements don't have any retail area described in OSM (including Weymouth, Wellingborough, Grantham, and Newark-on-Trent). 

The reasons vary.

  • In most of these the main retail area hasn't been marked as such, even though other types of urban landuse have been applied in other parts of the town. 
  • In some cases the whole town is marked as “landuse=residential”, 
  • In others an area that seems to be predominantly retail space has been described as “landuse=commercial”. 

Overall, there are too many exception cases to make productive use of retail landuse data.

On a very traditional high street, individual shop frontages tend to be relatively narrow, and relatively consistent in length. In theory, if we could compare the density of shops in OSM with what we expect, then we should be able to get a sense of where the database looks thin. However, this can only give a broad indication.


  • A well-documented street with a few large shops will still look empty in comparison to a street where the data only includes a small proportion of numerous small shops. 
  • There are practical difficulties in calculating density consistently when nearby shops may lie on opposite sides of a street, or in adjoining, and neighbouring streets. 

These challenges can be overcome to a degree, but my attempts have involved some intensive computation. In practice a simple heat map seems to be equally useful for flagging up suburban areas that already have fairly high levels of retail content. In conjunction with local knowledge of where retail outlets are clustered this could be sufficient to identify some larger suburban shopping areas that need attention.

This is an example from Newcastle. Data on retail is dense in the centre of the city, the quayside, and along Gosforth High Street. Coverage of retail barely shows up in areas such as Jesmond, and along Westgate Road. Those with local knowledge might be able to use this kind of feedback to identify areas that are worth further investigation.


Finally, while the overall level of retail coverage in England is 27% of retail premises, there are variations in the extent to which different types of shop are recorded. Analysis of the mix can be useful at a national level, but it can also be informative at a local level.

First, though, we will look at the mix at a national level.

Tuesday, 21 July 2015

OSM Retail Survey: Part-2

With 528,000 retail premises in England for a population of 53million, there is roughly one retail property for every 100 people across the whole country. It would be handy if we could use this ratio to examine coverage at a more detailed level than local authority.

The Office of National Statistics provides boundary data and population figures for Lower Layer Super Output Areas (LSOA) and Middle Layer Super Output Areas (MSOA). An LSOA has a population of 1,000 – 3,000 and an MSOA has a population of 5,000 – 15,000.

So we would expect to find 10-30 shops in an LSOA and 50-150 shops in an MSOA. However, when we use these ratios to measure actual coverage in OSM we find wide divergence. It is particularly noticeable that rural areas seem to be exceptionally well mapped, while suburban / residential areas appear under-mapped.



The underlying problem is that the number of retail premises is not proportional to the population at this level of detail. Suburban areas are well served by city centres, so have fewer shops than we expect. Rural areas with a dispersed population tend to have relatively large numbers of small shops - i.e. more than we expect. This diversity is demonstrated by examining how OSM coverage compares to the national average for different types of area. Sparse areas look well-mapped, even when they aren't. Urban areas don't look well mapped even when they are.

For what it's worth, at this level of detail, the correlation between the number of shops in OSM, and the number of residents employed in the retail sector  is even worse than the correlation between numbers of shops, and total population. So retail employment is likely to prove even less useful as an indicator of how many shops to expect, and I haven;t pursued this further.


ONS Rural / Urban classification
OSM retail units 
as % of expectation 
based on national average
Rural town and fringe
30%
Rural town and fringe in a sparse setting
83%
Rural village and dispersed
33%
Rural village and dispersed in a sparse setting
55%
Urban city and town
31%
Urban city and town in a sparse setting
61%
Urban major conurbation
32%
Urban minor conurbation
51%

As a result of these variations this approach is of limited use to us.

There are some variants that might be more useful. This example from North Tyneside highlights several Middle Layer Super Output Areas where there is no post-office recorded in OSM. Some of these might really have no post-office, but it's a fair bet that some of them really will contain a post-office, alongside other retail outlets that haven't been mapped yet.


So there may be some useful ways of using data from output areas based on population, but it turns out that it is probably more useful to examine the coverage of retail outlets across built-up areas.

Of the retail properties that I have found in OSM, 85% fall within a built-up area. It makes sense to look for numbers of retail properties within settlements. Again the Office of National Statistics provides us with handy geography and population data to work with. Here I'm using their data on population and boundaries of built-up areas to compare the volume of OSM retail data in larger towns and smaller cities. This is a less reliable, but a more granular view than we can extract from VOA statistics on retail properties that are available at local authority level.


This approach seems to work particularly well for mid-sized towns, and it can be adapted to make it useful for smaller towns and larger villages. We would expect most larger towns to be the main retail centre for the local population, so they should have roughly the average number of retail premises in proportion to the population. In practice, we find that well-mapped towns come close to this ratio. Where a town falls well short of the expected ratio it suggest that there is scope for improvement, and visual examination tends to confirm this impression.

We can see that Exeter and Chesterfield have roughly half the number of retail premises in OSM that we would expect to find on the ground. In Hull and Lincoln perhaps two-thirds of the retail premises are missing from the data.

Searches for retail locations will have higher utility to some people in some places than in others. If an application provider wants to use OSM retail data to support views of individual towns, then they might chose to begin with towns and cities of a manageable size, with large numbers of visitors, relatively high turnover in population, good technology infrastructure, etc.

University towns, for example, can be expected to have a high turnover of technologically adept students,



Cathedral cities are likely attract large numbers of visitors.



And so are Seaside towns.



Some towns fit into more than one of these categories, and quite a few of these look well-mapped (Bangor, Cambridge, Canterbury, Durham, Ely, Norwich, Oxford, Salisbury, Scarborough).  Others probably wouldn't take a huge effort to bring up to a similar level of coverage (Exeter, Worcester, York).

Measuring the ratio between retail premises and population begins to break down for smaller settlements. Retail is not evenly distributed, and we expect things to average out across a larger settlement, but not across a smaller settlement. The population of many smaller towns expect to travel elsewhere for some of their shopping. Some smaller settlements are predominantly residential. Others serve as retail centres for a wider area, so these have more than their fair share of shops and services. Similar variations apply in towns and villages that are popular visitor destinations. Nevertheless, it's unlikely that a settlement of 1,000 people would have only a couple of shops. It's not impossible that a town of more than 5,000 people will have no post office, or no pharmacy, but it seems unlikely. We should be able to use these assumptions (and others) to identify smaller settlements where retail is suspiciously under-represented in OSM.
  • There are 26 built-up areas with a population of more than 10,000, and 83 with a population of more than 5,000 where no Post-Office is recorded in OSM. The largest are Kirkby, Haverhill, Witham, and Formby.
  • There are 80 built-up areas with a population of more than 10,000; and more than 200 with a population of more than 5,000 where no Pharmacy is recorded in OSM. The largest are Braintree and Grantham.
  • There are 30 built-up areas with a population of more than 10,000; and 120 with a population of more than 5,000 where no food shop seems to be recorded in OSM.
I've used a mixture of these assumptions to identify smaller settlements close to home where the number of shops in OSM is implausibly low. It pointed me to a couple places locally that needed attention, and I have started to add shops. However, my home area is not a good example to illustrate the principle. Here, the under-mapped towns turn out to be quite widely dispersed, and don't show up well on a map. I'm less familiar with the locations in the more densely populated rural area around Durham, but it is a better example to illustrate the principle. The city itself is exceptionally well-mapped, but some of the surrounding small towns and villages look as though they might benefit from attention.



This approach of measuring the content across built-up areas seems more promising than using Output Areas, but it still has limitations. Individual settlements lying within a more rural landscape can easily be highlighted in this way, but it is more difficult to identify areas of incomplete retail mapping within a large conurbation.

To be continued....

OSM Retail Survey: Part-1

According to the Valuation Office Agency (the organisation which assesses business rates in England), there were 528,000 retail premises in England in 2012. This includes shops, banks, post offices,  cafés, restaurants, and take away food outlets, It doesn't include pubs or wine bars. It includes kiosks, some service providers (e.g. hairdressing salons), but it doesn't include filling stations.

To get a comparable figure from the OSM database we need to extract a mix of tags: most shops, some amenities, and some offices. So far I have found 142,803 features in England that fit the VOA categories of retail property (27% of the expected total). The shortfall across England is 385,000 retail premises that VOA have counted, but which I can't find in the OSM database.

The VOA statistics are broken down by local authority, so it is straightforward to compare OSM coverage at local authority level.


In notoriously well-mapped areas such as Nottingham and the neighbouring authorities of Broxtowe and Erewash the number of retail premises recorded in OSM is close to the number of retail premises reported by the VOA (Nottingham = 2,996 in OSM, 3,340 from VOA). In Tendring (Essex), the OSM tally is 95% of the VOA figure. In Oxford, and Cambridge it is almost 80% of the VOA figure.

I'm sure I must be missing some retail properties that I should be counting, and counting some that I shouldn't. Overall, though, these well-mapped areas suggest that I must be quite close to capturing what I hope to capture.

  • In 34 local authorities I reckon that more than half of retail premises are recorded in OSM,
  • In 24 local authorities less than one in ten of retail premises is recorded in OSM. 
  • The lowest levels of retail coverage are in Burnley, Castle Point, Doncaster, Eastbourne and St Helens.

Regionally, intense mapping around Nottingham means that the average coverage of shops in the East Midlands is relatively high (though not as high as the average across Inner London).

The lowest levels of coverage tend to be in larger northern towns and cities. Picking a few at random, I found less than one in ten shops recorded in Bolton, Doncaster, Rochdale, South Tyneside, and Sunderland. In the south, towns like Basildon and Luton do not fare much better.

Coverage of shops across rural counties is close to the national average.  The most complete rural areas border intensively mapped cities (Nottingham, Oxford, Cambridge). The least complete rural areas are widely scattered.

It doesn't really work this way, but imagine for a moment that the community collaborated to raise more areas to the standard achieved by leading examples such as Nottingham, Oxford and Cambridge. The obvious approach would be to prioritise areas that could be completed relatively easily, and the quickest wins would be where there is only a small shortfall between the number of retail premises reported by VOA and the number of retail premises recorded in OSM.

  • The smallest shortfalls between the VOA statistics and OSM content include authorities like Rutland (140), Maldon (176), Eden (201), Redditch (261), and Wokingham (269).

More generally, mapping of shops involves wandering around recording them. Comparing two authorities with a similar shortfall, it stands to reason that a more compact area would need less effort than the larger one:

  • There are just four authorities where coverage is already in the top decile, and size in the bottom decile, and all are inner London Boroughs – Camden, Westminster, Islington and the City of London.
  • There are nine more authorities that rank among the smallest 20% by area, and highest 20% by coverage. All are outside London. In addition to Oxford and Cambridge they include Cheltenham, Norwich, Redditch, Southampton, Tamworth, Woking and Worcester

By contrast, the local authority where I live covers a large area, has a low population density, and the proportion of retail properties recorded in OSM is below average. There is a long way to go before it ranks among the most thoroughly mapped. Random searching for shops could take a long time, so I need some way to prioritise.

To be continued....

Saturday, 15 November 2014

Bandstands

Bandstands are structures with a wide appeal, and the way they have been recorded in OSM throws a light on how contributors approach features that fall outside the mainstream. The OSM data on bandstands isn't complete, but the database may already contain one of the most comprehensive lists of existing bandstands in the UK.

The most extensive list of UK bandstands that I have found is the list of Vintage Bandstands, here. This has 334 distinct entries, but I don't think they all still exist.

In 2001 the Urban Parks Forum surveyed local authority parks in the UK. Out of the 438 bandstands they could identify 203 had already been lost, 186 were in use, or under repair, and 49 were abandoned or unused. Since then the Heritage Lottery Fund has been investing in the restoration and rebuilding of bandstands. I can't find more recent data, so for now let's assume that there are still more than 200 bandstands in public parks in the UK.

Bandstand in Sefton park, from Wikimedia

Across Britain, there are about 142 bandstands that are listed for architectural or historic interest. I have locations for those in England, and roughly three-quarters of them lie inside public parks, and roughly a quarter outside public parks. Around the same proportion of the bandstands in OSM are within a park, so it is probably fair to assume that by concentrating on parks, the survey by the Urban Parks Forum was primarily concerned with about three-quarters of the bandstands in the UK.

So as a starting point I'm going to assume that there are about 275 bandstands left in Britain, of which about 213 are inside public parks, and 62 outside public parks - fewer than on the Vintage Bandstands list, and more than the Urban Parks forum suggests.

I've managed to find 223 bandstands in the OSM data for Britain, which is about 80% of my estimated total. If these features really are bandstands, then that's an impressive result for a feature that I thought would fall well outside the mainstream.

Our ancestors obviously knew how to build structures that would maintain their appeal, and that's part of the attraction of examining data on bandstands. But I also wanted to look closer at the data because the tagging of bandstands in OSM is particularly inconsistent, and I thought we might learn something from that.

A fairly simple search of the OSM database for anything that mentions "bandstand" (and spelling variations) will pick up 247 features. On inspecting the data we find about dozen of these are false positives: completely different features with the word "bandstand" in the name. Ten appear to be duplicates. If two different features within a few hundred yards of each other both describe a bandstand then they are probably referring to the same structure in the real world. Some of the duplicates occur because contributors have added and tagged both a way and a node. Some may be because one attempt hasn't rendered and a later contributor thought the feature was missing. In a few cases contributors seem to have been uncertain how to tag the feature, so they have added more than one option.

About 50 of the bandstands in the OSM database correspond to the 93 listed bandstands in England, so contributors have added more than half of the bandstands that are listed by English Heritage. Very few have been marked as a structure with listed building protection. If the overall totals are correct, then contributors have located more than 80% of all bandstands, but less than 60% of listed bandstands. I'm not sure what to make of that.

For processing this data we really want to find bandstands based on well-defined attributes. Interestingly, with bandstands we find more variation in the choice of keys than in the choice of values. The most common contents of the value, by far, is "bandstand", with "band_stand" accounting for about one in thirty values. The keys that have been marked as a bandstand include "leisure", "amenity", "building", "historic", "type", "building:use", "tourism", "shelter_type", and "man_made". It's interesting that bandstand contributors choose quite a wide variety of different keys with quite a narrow range of values.

There are 233 features that are recognisable as a bandstand from the data content. This includes 10 that appear to be duplicated, which mucks the numbers up slightly. I've slightly fudged this in the chart - to keep things simple.



The tag "leisure=bandstand" is recommended in the documentation, and accounts for about 24% of all UK bandstands (30% of the ones I found, and almost half of the bandstands that have been coded in a structured way). The "amenity" and "building" tag with a value of bandstand account for another 18% of all bandstands. Less common tags such as "historic", "type", "building:use", "tourism", "shelter_type", and "man_made" account for another 2%. Low usage of "man_made" surprised me, because I thought this would have been seen as more appropriate than "building". Apparently not.

More than a third of bandstands can be identified in the database by the name, but not by other coding. This approach, of course is risky for automated processing, because the results need to be checked for false positives. However, for some purposes it is still worth looking at how bandstands are named. The form "name=bandstand" is the most common - almost as though the name tag is being used for coding, since this is not capitalised. Less common forms are "name=Band Stand", "name=The Bandstand" and "name=The Band Stand". I suspect contributors have been influenced here by labeling in the standard render.  Together, these four values pick up 84% of the named bandstands. The rest mainly use names based on the location - such as "Southsea Bandstand".

The few remaining examples that I found are a mix of more obscure tagging, and spelling variations. These are of little interest, or value to data users.

There are probably about 50 UK bandstands that don't appear in the database (yet).

We have to be wary of false synonyms. I've looked at features in the database that correspond to the location of listed bandstands, and along with tagging variations in the OSM data these suggest that some contributors consider "gazebo" and "pavilion" as synonyms for "bandstand". Some just label the feature as a "shelter". A "gazebo" can look similar to a bandstand, although a bandstand is generally larger, raised higher above ground level, and clearly intended for a different purpose. The term "pavilion" might be acceptable as a technical description of the architecture, but in general data users will not be able to use it because it is so widely used to mark a sports pavilion. And "shelter" is a very general category, that doesn't help somebody who wants to identify bandstands.

Now to draw some conclusions.

Although we know for certain that some are missing, there's a case that the OSM database already contains one of the most complete lists of existing bandstands in the UK. With relatively little effort to plug the gaps, and verify existing data this is information that could be used productively, by anyone who wants to do so.

If somebody wants to find bandstands in the database at present they will probably look for values of "leisure", "amenity" or "building" that contain "bandstand". That will uncover almost half of all examples in the database. It starts to get quite complicated for data users to seek out tagging variations and find the next few percent. Searching the "name" tag for variants of "bandstand" will turn up quite a few likely candidates, and might be appropriate in some circumstances,but it would not be reliable enough for systematic processing. Without manual inspection this approach also picks up theatres, cafes, and the like that have been named this way.

The data on bandstands is fairly comprehensive and suggests that contributors can have quite different perspectives on these features. So this is an area where we could (and probably should), encourage more consistency, while tolerating quite a lot of variation in tagging. For the relatively small number of features in the database, bandstands demonstrate an unusually varied use of tags. Whether they realise it or not, different contributors have been marking bandstands according to their function ("leisure=bandstand" or "amenity=bandstand"), according to their form ("building=bandstand", "man_made=bandstand"), or according to their significance ("historic=bandstand"). These are surely all valid approaches, and they are not incompatible with each other. Indeed, quite often they are used together on the same feature. It is quite conceivable that one bandstand will be notable for its historic significance, but no longer in use as an amenity, while another might be in regular use as a leisure facility, but have no historic significance. Some might have changed function ("shelter:type=bandstand"). We don't know how this data might be used in future, so differences such as these ought to be reflected somehow in the tagging. However, the current documentation doesn't give guidance on such subtleties, so the current data probably doesn't record them accurately. All we can really say for now is that features with any of these tags have been recognised and recorded as bandstands.

If the community wants to improve the current data on bandstands then the following might be priorities :

  • locate the fifty or so missing examples
  • add structured tagging to the hundred or so bandstands which can currently only be found by name
  • fix inconsistencies - such as combining duplicates
  • encourage considered use of the existing tagging options to capture and retain information that can be collected in the field
  • add attributes (such as listed building status) that could be of interest or value to data users
  • find a group of bandstand enthusiasts who might be interested in verifying the data, finding innovative uses for it, and taking things further
  • celebrate with an outdoor concert

Tuesday, 11 November 2014

Kennels and catteries

Kennels and catteries are commercial businesses where owners can leave their cats or dogs while they are away from home. They are part of a wider category of animal boarding that also includes donkey sanctuaries, and organisations that take in domestic strays, pets whose owners are no longer able to cope; and wildlife, such as hedgehogs. These are more likely to be run by charities.

In England and Wales animal boarding establishments (including kennels & catteries) are controlled by the Animal Boarding Establishments Act 1963. They have to be licensed by the local authority. The situation is similar in Scotland, but controlled by a different act.

Because they have to be licensed I thought it would be straightforward to find statistics on how many kennels and catteries there are in the UK. I was mistaken. Published government figures on the business population don't go down to that level of detail, local authorities don't seem to publish any statistics, and I can't find figures from trade bodies. However, the Valuation Office Agency does publish figures for different types of property, including the numbers of kennels and catteries there are in Wales and the English regions. These figures are a bit old (2010), but broadly in line with the numbers that come up on a search of Yellow Pages. Unless anyone can come up with a better figure, I think we can be fairly confident that there are just short of 5,000 kennels and catteries in the UK.

Of these I've been able to find just over 200 in the OSM database. Contributors have tagged roughly half of those as a kennel or cattery, and the other half can be identified (with reasonable confidence) as a kennel or cattery by the name.



Mapping these clearly hasn't been a high priority for OSM contributors.

I reckon that the data on kennels and catteries is too incomplete, and the tagging is too inconsistent for it to be of great practical use for rendering or other forms of data presentation (at present). The point of this post isn't to argue that things should be any different. These establishments are not particularly prominent features in the landscape. Dog and cat owners (in my experience, at least) will either have chosen their a preferred animal boarding service already, or they will find one through personal recommendation rather than searching a database. This isn't quite the same for animal rescue, where I could see a need for an application to find the nearest hedgehog sanctuary (for example). But it's hard to see how pressure from data users is going to create a surge of interest in data on animal boarding. One day we may see enthusiasts kick of an "Animal Boarding Mapping Project". But my guess is that these are more likely to be added by non-specialist contributors who are working on mapping a wide range of different features within their local area.

In any case, the community will decide on priorities. My interest in looking at this is not to push for action on kennels and catteries. I'm more interested in seeing what we can learn about how contributors approach less commonplace features.

We find three different models for tagging kennels and catteries.
  • The usual approach is not to label these as a kennel or cattery at all. Examples of kennels and catteries can be  picked up fairly easily through the contents of the "name" tag. There may be other, similar techniques that I haven't tried. In other words these features have been mapped, and named, but there is no further detail to indicate that they might be of special interest. In the chart they appear as "Name only"
  • The second most common approach is to use simple "amenity" tagging. This makes use of user-specified values of the "amenity" tag: "amenity=kennels" and "amenity=cattery". These aren't documented, but nevertheless, they represent more than half of the examples of kennels and and catteries that have appropriate tags attached in the database. In the chart these appear as "Simple"
  • The third approach is more structured, and follows the tagging recommended in the documentation. This is based around "amenity=animal_boarding"  with more detail added under "animal_boarding=...". This approach represents almost half of the examples of kennels and catteries that carry specific tagging. Although it is the documented approach it is not yet the most widely used. It appears in the chart as "Structured".
As always, there are a few variants on both of the structured models. By the look of it most of these are a mix of typing mistakes, and misunderstandings. They don't make a significant difference to the totals. There are an even smaller number of intriguing examples, though, where a contributor has used tagging based on "pet=". There are only a few of these, and it looks as though this might be an experiment that didn't go further, but it suggests that at least one contributor sees boarding kennels and catteries in terms of how they relate to other facilities for pets, rather than as a type of amenity, or as a service for animals.

The blindingly obvious conclusion is that there are serious limitations in the current data. Inconsistent tagging might be a deterrent for users who want to render or process this data: but the real blocker is the very limited coverage.

In the case of kennels and catteries, it seems likely, as things stand, that anyone using this data is unlikely to be interested in rendering or presenting the data within an application.Simply because there isn't (yet) enough coverage to make this viable.

There is no shortage of similar examples in the database, outside the mainstream,  where coverage is low.

To my mind this raises some interesting questions.

When the community is debating about best how to tag features that fall outside the mainstream:

  • Who is using this data, and what are they using it for?
  • Do current approaches meet the needs of data users, can they be improved, and if so, how?
  • Could contributors be encouraged to add more useful data?

Presumably we think that one day coverage of these features will become sufficient to make rendering or application processing viable. If we don't think that, then why are we collecting this data at all?.

  • Will the needs of data users change at that point?
  • If needs do change, how will that affect the data?
  • How will we tell when we have reached that stage?
Obviously I wouldn't be raising these questions if I held the view that the needs of data users were the same for all types of feature, and that they were unchanging as the contents of the database evolves. 

To me, the key question is "how do we tell when we have reached the stage that rendering or data presentation becomes viable?". But this might touch on some contentious issues. It's probably best to stop at that point, and see what others think.

Saturday, 8 November 2014

Bookies

The Gambling Commission publishes statistics on the number of bookmakers in Britain. Earlier this year they recorded 9,021 bookmaker premises, based on returns from operators. The number has been fairly constant in recent years.

A data user will be able to find about 1,280 of these in OSM (depending on how determined they are).




The preferred tag is "shop=bookmaker". That picks up 8% of all bookmakers in Britain. The most common alternative is "shop=betting" and that will pick up about 4% of the total.

There are a number of variants on these (shop=bet, bookmakers, betting_shop, bookies, turf_accountant) Together these pick up about 0.4% of the total.

A few contributors have gone down slightly different routes. Using the "amenity" tag rather than the "shop" tag accounts for another 0.2%, and various bookmaker-related values for "gambling" account for another 0.4%.

There are also some premises which look like a bookmaker (based on the operator or the name), but aren't marked as such. I found another 79 premises that might have added to the collection if they had been tagged with additional data. There are almost certainly more that I missed, but deep searches to find these get increasingly complicated and the results increasingly suspect.

I'm unable to find around 86% of British bookmakers in the OSM data.