Hello everyone.
I quickly built this little tool based on public records provided by the government.
The amount of information they provide is incredible. There are no explicit names, but for small startup or exec position it is easy to guess who is who.
As far as I know this is the biggest LCA database so far.
Thanks for this! I'm currently looking at opportunities and this would help immensely during negotiations.
Would it be possible to search for role + location instead of company first? Essentially the refine functionality you provide on company's yearly page [0] to be at the root. You can perhaps make state or city mandatory so that it doesn't hit your backend hard.
Great implementation. I've seen the raw data before but never had the web development skills to execute on this idea. Can you also add a query where once you select a company and year, it groups by title and gives average salary.
I actually think it means they decide to no make go further with the LCA. Like they decided to finally not hire the employee or the employee decided to not work for this company anymore.
Actually that makes more sense. It is possible to be working on another visa while a company files the LCA for you (TN, L, visas, etc), but it's probably less likely than a new H1-B hire.
For a company like Netflix, my guess is, you need to be a world-class expert at some area of scalable backends and/or big data. For example, you need to be a core contributor (in terms of concepts, architecture etc. - not just coding) to projects like Kafka, Cassandra, Spark etc.
Very impressive. I'm organizing an Immigration-themed "Startup Weekend" event[1] happening next week (in SF), and I think I'll be sharing this with participants there as an example of what can be done.
Instead of plotting a count per exact salary could you plot count per binned salary? Ie instead of plotting 1 at 105k and 1 at 107k plot 2 at 100-110k.
As it is the plots aren't particularly useful for comparing the distribution of salaries between companies etc.
Yes I am still cleaning the data. It's all unbelievable the amount of crap you can find in the public records. I have a lot of duplicates unfortunately.
Down the road, I'm planning on providing a couple of cool charts/map about trends and evolutions.
- ratio of LCAs to total permanent full time employees. This will show which companies are really leaning on the H1-B visa (I'd estimate my "household name" software company is at about 50%)
- Source of prevailing wage. Interestingly employers don't have to use the BOL published data and can self report. I'd be interested to see how many self report.
- The average delta between prevailing wage and salary per employer/job title.
Check out Open Refine. Has a feature that clusters similar strings and unifies. I remember last time I looked at this data set... 4 letter acronyms spelled 12 different ways, it's unbelievably messy.
The amount of information they provide is incredible. There are no explicit names, but for small startup or exec position it is easy to guess who is who.
As far as I know this is the biggest LCA database so far.