Friday, November 21, 2014

Policy: The Software Patent System Is Broken

(Our Broken Software Patent System, Part 1 of 3. Or you can read the short version)

The American patent system is completely broken, at least when it comes to software patents. Billion-dollar software companies are suing each other over minor pieces of code, patent trolls run rampant, and consumers are losing out as the original purpose of software patents have been lost. In this three-part article, I'll first convince you that the software patent system is broken in the US, then that the system shows no signs of changing, and finally I'll discuss what we can do about it.

Note that this article is about software patents in particular. The US patent system may work better for other industries; I don't know. My research has been confined mostly to software patents.

The Point of Software Patents

Let's first talk about the point of software patents. Patents are a form of intellectual property, like copyright. Unlike copyright, patents cover inventions instead of creative works and last for a shorter time (20 years compared to copyright's roughly 100 years). The purpose of a patent is laid down in the constitutional clause that gives Congress the power "[t]o promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries" (Article I, Section 8, Clause 8).

The idea behind patents is great. If you spend years inventing something, perhaps building a better mousetrap, then someone shouldn't be able to see your mousetrap in a store, steal your idea and sell the same mousetrap while bypassing all the hard work you did refining the improved mousetrap.  Or worse, they could see your mousetrap in your secret lab while you are still refining it, steal your idea and beat you to market.  With patents, you must reveal your invention publicly, but then you have the exclusive rights to that invention for 20 years.

However, there are limits. You can't patent just anything, someone would patent the wheel and sue all the car manufacturers.  Patents must be a patentable thing (a process, machine, "[article] of manufacture," and composition of matter), and they must be  new, useful and nonobvious. "New" means you can't patent the wheel; it's already been invented. "Useful" means the invention has to actually do something. "Nonobvious" means you can't just slightly tweak someone else's invention; if there's already a patent on a standard pencil, the patent office won't grant you a patent on a slightly smaller pencil.  That's just an obvious extension of an existing invention.

Applying these principles to software is pretty new. Software patents only became legal in the US in the 70's and 80's. The courts are still trying to figure it out. In eBay Inc. v. MercExchange, L.L.C. (2006), Justice Kennedy and other justices questioned the wisdom of permitting injunctions in support of "the burgeoning number of patents over business methods," because of their "potential vagueness and suspect validity" in some cases.

Lots of Existing Patents Are Invalid

Non-New Patents
Now that we've discussed how software patents are supposed to work, what's the problem? Well, one problem is that people are patenting things they shouldn't be able to patent. Remember that all patents must be new, useful and nonobvious.  Well, people are patenting existing software inventions.

Two excellent sources of information about our broken software patent system are the This American Life episodes " When Patents Attack!" and " When Patents Attack ...Part Two". I cannot recommend these episodes of This American Life highly enough. You can listen to the first one below. But to get to my point, during Part One the journalists from This American Life (TAL) talked to David Martin, the founder of M-CAM. M-CAM is hired by governments, banks, and other businesses to assess patent quality. M-CAM uses special software to search through all existing software patents to find patents that are essentially the same. Ideally, each patent should be unique. However, sometimes the US Patent and Trademark Office (USPTO) grants patents for things that have already been invented and patented. During the TAL episode, Martin looked up how many matching patents there were for one specific software patent. There were 5,303 matching patents.
"We thought that would be an anomaly. And then we were told, oh no, it's not an anomaly. That happens. [....]And as I've testified in Congress, that happens about 30% of the time in US patents."
- David Martin, (emphasis added)
The specific software patent that Martin was looking up with TAL was Patent 5771354 "Internet online backup system provides remote storage for customers using IDs and passwords which were interactively established when signing up for backup services". According to the owners, Intellectual Ventures, it covers upgrading software on your home computer over the Internet. For example, "when you turn on your computer and a little box pops up and says, 'Click here to upgrade to the newest version of iTunes'". When TAL looked at the text of the patent it seemed like it did a lot more than that. Having looked at it myself, I'd agree. It appears to cover doing any sort of data backup over the Internet, using an ID and password.

Drawing from Patent 5771354 showing
how all computer networks work
Didn't online backup solutions exist before 1999, when this patent was filed? Apparently. The Wikipedia article for cloud storage talks about AT&T's PersonaLink Services in 1994. This was a service for storing your PDA data. That sounds like an existing invention to me. But let's ignore that example for now. Martin's software found Patent 6003044 and Patent 5933653, which also cover backing up data remotely.  This invention has been patented many times over.  When TAL's Alex Blumberg expressed incredulity that 30% of US patents were redundant, Martin pointed out there was a patent on toast.

Patent 6080436 "Bread refreshing method" was issued in 2000.  It patents toast.  Don't believe me?  Go read it for yourself.  Granted, it isn't a software patent, but it shows the flaws of our patent system.

It should be clear now that the USPTO allows inventors to patent inventions that have already been invented, breaking the requirement that all patents be new.  There are duplicate patents abound.  Furthermore, I'd argue that even if an inventor receives the patent on something, it doesn't mean he or she invented it.  That person may have just been the first to file a patent.  Filing a patent is an expensive, time-consuming process.  According to the USPTO, it takes about 2 years for a patent to be processed.  The cost of filing a software patent is around $10,000 (UpCounsel, Richards Patent Law, IPWatchdog) when including legal fees, which might be a drop in the bucket for a major corporation but is expensive for a small startup.  Because of these barriers to getting a software patent, the actual inventor may not be the first to file a patent on a particular invention.

Obvious Patents
So software patents may be patented and re-patented.  But even "original", previously unseen inventions may have problems.  For example, they might be obvious extensions of existing patents.  As I said above, you shouldn't be able to patent a slightly smaller pencil.  The pencil has already been invented, so taking an obvious idea such as making slight tweaks to the size or color is not patentable.  Sure, that idea might be "new", but inventions must also be nonobvious.  But obvious software inventions are being patented every day.

As a software engineer and a person who is skilled in the field, I would describe Amazon's 1-Click patent as fairly obvious.  In short, it describes the ability to save your credit card information in a website so that you can purchase items with only one mouseclick.  I am not the only who believes this patent to be obvious.  As one blog pointed out, it's a "fairly broad concept" that the European Patent Office denied a patent to because it was "obvious to a skilled person".  To give another example, this patent caused the Free Software Foundation to boycott Amazon for a short time on what it called "an important and obvious idea for E-commerce".

There are other examples of patents on obvious "inventions."  Richard Stallman, a software engineer who is famous for his work on GNU, Emacs, and other free software, decries the obviousness of Patent 5963916, applied for in October 1996 in his text, "The Anatomy of a Trivial Patent".  Patent 5963916 ("Network apparatus and method for preview of music products and compilation of market data") covers listening to a preview clip of music on the Internet.  After posting a snippet of the patent, Stallman writes, "That sure looks like a complex system, right? Surely it took a real clever guy to think of this? No, but it took cleverness to make it seem so complex. Let's analyze where the complexity comes from[...]"  Stallman then proceeds to analyze each line in the first portion of the patent to point out how each aspect of the idea was already existing or, at least, very obvious at the time:
Now look at a subsequent claim: 
3. The method of claim 1 wherein the central memory device comprises a plurality of compact disc-read only memory (CD-ROMs). 
What they are saying here is, "Even if you don't think that claim 1 is really an invention, using CD-ROMs to store the data makes it an invention for sure. An average system designer would never have thought of storing data on a CD."
In case it's not clear, Stallman is being extremely sarcastic. Even back in 1996, CDs were commonly used for storing data.  This is one of the patents that Stallman calls "laughably obvious".  The problem is that the overly complex language of patents obscures the obviousness of the ideas.

This is a huge problem.  Lots of software patents in the US are actually invalid because they are obvious or already patented.  If David Martin's numbers are to be believed, there are hundreds of thousands of invalid software patents in the system.  But invalidating a patent means spending lots of time (years) and money (millions of dollars) in court.  Because of this, the original purpose of software patents--promoting progress--has been lost.

Software Patents Aren't Serving Their Purpose

The first problem with software patents is the barrier to entry.  That is, the money and time spent in getting a patent may create a burden for startups and solo software engineers.  As mentioned above, it can take around 2 years and $10,000 to get a software patent.  This fact alone may dissuade programmers from trying to create innovative new software that the public would benefit from.  However, this is a small problem compared with other software patent issues.

One of the most insidious problems with software patents is the menace of patent trolls.

Patent Trolls
Alaska Robotics - "Patent Trolls"
According to the most popular definition, a patent troll (also called a patent assertion entity) is a company that obtains patents--usually through buying them--and, instead of making products based on those patents, waits for another company to violate their patents, and then forces that company to pay licensing fees to use that patent.

Let's take the example used in the introduction of the "When Patents Attack!" This American Life podcast.  Since 1999, Jeff Kelling has been working for FotoTime, a small company that hosts a photo sharing website.  This was before major photo sharing sites like Flickr, although FotoTime wasn't the first photo-sharing website.  In May 2008, they received a letter from FotoMedia (not to be confused with FotoTime) that said FotoTime was in violation of 3 patents.  FotoTime was told to contact FotoMedia to arrange payment, or FotoMedia would take them to court.  Jeff's team looked up the lawsuit and noticed that FotoMedia had sent the letter to 130 other companies including Yahoo (Flickr), Shutterfly, Photobucket, and other companies, big and small.

There were a lot of reasons this was weird.  One was that FotoTime hadn't realized that they were violating any patents.  Whatever patents were being violated, FotoTime had evidently come up with the same idea themselves, without reading FotoMedia's patent or copying some FotoMedia product.  Another weird thing was that FotoMedia wasn't a competitor to FotoTime.  FotoMedia didn't have a photo sharing website.  Then when Kelling called FotoMedia to ask them which patents FotoTime was violating, FotoMedia wouldn't tell them.  Kelling said, "They said they wouldn't answer that until we got into court."

Kelling learned that fighting the patent would cost an estimated $2 to $5 million, which was "more than [FotoTime] could handle".  FotoTime settled with FotoMedia.  As part of the settlement, FotoTime is forbidden from saying how much they had to pay FotoMedia.

This is what patent trolls do.  They purchase a patent that is somewhat vague or perhaps even invalid, and instead of making some technology that uses the patent, they threaten to sue companies they think are using the patent.  They might try to sue hundreds of companies, as in the case of FotoMedia.  Some companies will pay them a fee rather than spend millions of dollars fighting the patent in court.  The patent trolls will then use that money to purchase more patents and sue more people, ad infinitum.

This isn't a business.  This is, as venture capitalist Chris Sacca described such practices on This American Life, "a mafia-style shakedown."

And patent trolls are growing bolder.  According to a Presidential study, the number of lawsuits brought by patent trolls more than doubled from 2010 to 2012, and these lawsuits accounted for 62% of all patent lawsuits in America in 2012.  Victims of patent trolls paid $29 billion in 2011.  Now in 2014, these numbers are probably higher.

The Damage from Patent Trolls
This patent trolling kills innovation in a few ways.  One problem is that software companies must spend money on legal fees instead of growing their technology.  It's impossible to know whether your company will be violating a patent or whether a patent troll will target your company, so it's now important to have a good legal team to either negotiate a settlement or fight the patent lawsuit if and when a troll sues your company.

Some startups may not be able to pay the legal fees from a patent lawsuit and may go under.  Jeff Kelling of FotoTime said that "The settlement they [FotoMedia] wanted to get was just enough to put us in danger, but not to close us [...]"  It seems that lawsuit came close to shutting down FotoTime.

Patent trolling can even create a chilling effect, causing programmers good ideas to never pursue their startup for fear of being sued by a patent troll.

Software patents no longer promote progress, if they ever did.  Now, software patents stifle innovation.  It's impossible to know whether a company is actually violating a patent because the patent office is a confusing mess, containing vague and duplicate patents.  And so any software company with a good idea can be sued at any time, by a patent troll claiming that software company is violating its patents.  Sure, the current software patent system is great for patent trolls, who create nothing and provide no benefit to society.  And this is the exact opposite of the purpose of the software patent system.

The system needs fixing.

Continued in The Software Patent System Will Always Be Broken

This American Life: When Patents Attack!

Saturday, August 2, 2014

Software Engineering: Good Reads

I'm not an avid book reader, so I feel that my book suggestions should be taken seriously, since there are so few books I can really recommend.  I think there are a few software-engineering-related books that every good programmer should have on their shelf.  I will list these books, roughly in order of importance.  I've also listed each book's latest edition (that I know of).

The C Programming Language, 2nd Edition by Brian W. Kernighan and Dennis M. Ritchie

The C bible. C is still an important language for many programmers to learn, and this is the book to use to learn it.  Never before or since has a programming language had a textbook so thorough.  I only wish that Java or C++ had a similar textbook that I could recommend as highly.  Maybe it's because C is a small language, but this book works well whether you're learning C for the first time, or you just need a reference.  The only downside is that the latest version of this book covers ANSI C, instead of either the newer revisions, C99 and C11.  For that reason, if you're using a C99 or C11 compiler, this book is somewhat deprecated. C99 made a lot of important changes to the language.  However, the book is still a good starting point if you don't know any C (or, again, if you're using an older compiler).

Introduction to Algorithms, 3rd Edition by Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, Clifford Stein

This book contains all the information you need to know about algorithms and data structures, including their pseudocode implementations and computational complexity.  I suggest it to everyone who is applying to Google as a software engineer.  The only issue is that it is too complete, and contains a lot of extraneous information.  It's not a book you read cover to cover.  Also, I'm not a big fan of their "pseudocode."  A more Java-like or Python-like language would be easier to read.  As a software engineer for any company that relies on speed for computing large amounts of data, you will need to know the most efficient way to solve your problem.  This book will help you get there.

Effective Java by Joshua Block

That's great that you can write a one-line method to sort an array.  Now, graduate from a programmer to a software engineer by learning why that is awful.  Effective Java tells readers about readability and other important issues that come up with writing code in a work environment.  Basically, it's a collection of tips for writing readable, maintainable Java code.  Do you swallow InterruptedExceptions, that is, catch them and then simply throw them away?  If so, Joshua Block will tell you why you're breaking his heart and the heart of everyone around you.  (For those who just want the answer, read this. Or the simple answer is that you may prevent a program from being killed when it needs to die.)

Software Project Survival Guide by Steve McConnell

This book is written more for managers and team leads, but is useful to all engineers.  It describes in detail the basic ideas and reasoning behind a waterfall software engineering process.  A lot of you may be using agile or scrum, but you can still incorporate most of the ideas in this book in your work process.  Testing, maintainability, requirements, tracking... this book has all you need to turn your cowboy coders into seasoned engineers.  Process, process, process.

Team Geek by Brian W. Fitzpatrick and Ben Collins-Sussman

This book written by two Google engineers seems to be written more for team leads, but, again, is useful to any software engineer or manager of software engineers.  It's all about the non-technical side of software engineering; the human element.  For example, they cover dealing with "poisonous" coworkers, building a strong team, how to have efficient meetings, and how to communicate and collaborate well.  This is the only book I didn't have to read for school or work that I'm recommending.  I think it's important because even a group of smart, trained software engineers can have interpersonal issues or leadership issues, and this book tries to train people to alleviate those issues.

Friday, June 27, 2014

Software Engineering: Automated Software Testing 101

My official title at Google is Software Engineer in Test Tools and Infrastructure, which means I spend most of my time writing automated test infrastructure and consulting with teams to help them test their products.  For the last couple years that I've been working in this role I've learned a lot about testing.  I'd like to disseminate some of this knowledge.

There are many ways to test your software--way too many to cover in one short blog post--, but some of them are more important and widely applicable than others.  In this article I'll talk about the major types of automated software testing and why they are important.

Why and When to Test
Firstly, why do we need to test software?  Software testing can sometimes seem less important than other tasks.  For example, is there any reason to spend expensive engineer-hours writing tests, when you already have design reviews, code reviews, and manual code inspection to assure product quality?  Additionally, the fact that bugs can still exist in well-tested code is disheartening.

Despite the drawbacks, testing should be done early and often.  Testing your software increases product quality (and in doing so, lowers maintenance costs), aids development speed (testing takes time, but bugs are easier to fix when found earlier), proves requirements have been met, and can reveal design flaws.  It's usually worth the time to write at least a few small tests.
"The problem with quick and dirty, as some people have said, is that dirty remains long after quick has been forgotten."
— Steve McConnell, Software Project Survival Guide
(By the way, the Software Project Survival Guide is a pretty good book for those in the software field. I recommend it for engineers and especially for project managers).

This article focuses on automated tests because it's my area of expertise, manual tests are pretty easy to figure out, and because automated tests have several advantages over manual tests: automated tests are generally faster to run, can be run more frequently when part of a continuous build, are easier to use for unit and integration testing, and free up engineers to do other things.  If you're writing a small program for yourself, relying solely on manual system tests is fine.  For major projects, you'll need automated tests.

What to Test
So by now you've been persuaded by my convincing words that testing should be done early and often. But what are we testing?  If we test everything, at what point do we write tests for tests?

Tests should be created for most things, especially complex and critical functionality.  If you've created a function that does nothing but add a digit to the end of a string, it's not critical to write a test for that function. Most other stuff should be tested.

While tests themselves can break, tests don't usually need their own tests because of their simplicity and because a false failure will be visible to your team when the test runs.  A falsely-passing test can be a disaster, but is an infrequent occurrence for well-written tests that were working to begin with.  The exception is to write tests for test infrastructure.  Sufficiently complex tests are usually built on some sort of testing infrastructure (JUnit, WebDriver, a fake database, etc.) that can easily contain bugs, and accordingly, they should be tested.

Who Does the Testing?
Who writes tests: software engineers or a QA/test team?  The answer is that it depends.  For small teams creating small projects, it is sufficient to have engineers test their own code.  After all, it can be difficult to obtain a test guy or gal, and there isn't much code to test, so writing tests is quick and easy.  But even for large teams creating large projects, it can be beneficial to have the programmers write tests.  They can gain insight into their own code, have a better idea of what can break, and are more familiar what the issue is when a test breaks.  As someone who writes tests for other teams' projects, I've often encountered teams that have no idea why a test is broken or how to fix it, even though they know the project code better than I do.  That issue diminishes when the team has a hand in writing the tests.

Alternatively, a separate test team gains a different perspective of the code.  Unlike the programmers that wrote the code, a test person's ego is not affected when a test finds a bug.  In fact, finding bugs is ego-boosting for the test team.  The drawback is that a test team may not write tests that sufficiently cover the weak points of the dev team's code.

So the answer is: either way is pretty good.  Get a test team when writing tests becomes burdensome for the project programmers.  This can happen when the software is extremely hard to test for whatever reason, or when it's complex enough to require special test infrastructure.

Test-Driven Development
So tests should be written early.  But how early?  Some programmers believe in test-driven development, where you write the test before writing the code.  I think that's a bit drastic, but not always a bad idea.  APIs and features may change slightly when actually writing code, so your tests will likely have to change anyway.  You don't necessarily need to write a test before the code is written, but you should have a test plan.  Writing tests should be done simultaneously when writing code; that's early enough to catch initial bugs and late enough to prevent overhauling tests.

Now, on to types of tests...

Unit Testing
First and foremost, if you're going to have any automated tests (i.e. you're not writing a small program for yourself), you should have unit tests.  The reason is that unit tests are the easy and quick to write and run, unlike other tests.  Unit tests are small tests, usually white-box style, that test a class, a few classes, or another small portion of code.  Using mocks or stubs is perfectly fine for unit tests.

Because unit tests are so lightweight and quick to run, they should be run often, preferably as part of a continuous integration process.  This allows the unit tests to catch newly introduced bugs quickly.  Several solutions for continuous integration exist, like Jenkins. (I have not used Jenkins myself).

Integration Testing
The downside of unit testing is that important parts like databases or dependencies on other binaries* are mocked out or nonexistent.  You don't get a good picture of how the software performs under real conditions.  In order to verify the different subsystems of a program, you should write integration tests.  Integration tests are large tests that test multiple binaries, or one binary that uses external resources.  Integration tests are helpful in catching integration errors (often API or design bugs).  Even if two pieces of a project are "bug-free," they may not integrate well, resulting in miscommunication that can only be caught by integration tests.  Integration tests can also find speed or memory issues that can't be found by simple unit tests.

Since integration tests are large and slow, it can be difficult to find a test framework that will bring up all the resources (a.k.a the environment) needed for the test and then tear them down after the test is done.  One workaround is to leave up servers or other external resources indefinitely and have the tests clean up any permanent effects in the environment after the tests finish.  This is a dangerous game, as the long-running environment can get into a bad state, giving inaccurate test results.  Sometimes the only way to run integration tests is to manually bring up and down the environment.  In this case, it might be worth it to manually run the integration tests as well.

System Testing
To get a complete picture of how the software will actually perform, system tests are required.  System tests are large tests, black-box style, that bringing up an actual environment and simulating an end-user using your product.  Not only do they help reveal memory and speed issues that may only occur in a real-world environment, they also may reveal UI and usability issues.

System tests are often run manually since it's hard to find good frameworks that can both setup a large environment and give human-like input (often by manipulating a GUI).  However, many GUI-manipulating tools exist.  I've used AutoIt to manipulate Windows GUIs with excellent results and I've also used WebDriver to manipulate webpages with very good results.  Combined with integration-test-style frameworks, you can potentially automate your system tests.

In summation:
  • Write tests early and often
  • Run unit, integration and system tests, automating them if you can
  • Run automated tests as part of a continuous integration process if you can

Following these suggestions will help you get excellent code coverage and excellent feature coverage, which will result in stable, easy-to-maintain, and on-schedule software.

Update: Added What and Who testing sections.

*Binaries are individually-compiled programs. They may not do much alone, but work in conjunction with other programs to create useful output.  If they do useful work by themselves, they are standalone binaries, or executables.

Tuesday, May 20, 2014

Policy: Open Letter to the FCC

The FCC has established an email address (openinternet@fcc.gov) for public comments on their Open Internet proposal.  I sent the letter below.  All comments become official FCC proceeding documents, so they are publicly available via the web.



FCC Commissioners,

My name is Omari Christian and I'm a software engineer.  I have read and written a lot about net neutrality and I am very concerned about Tom Wheeler's new proposed rules for broadband.  Particularly troublesome is giving Internet service providers (ISPs) leeway to charge content providers in a manner that is "commercially reasonable."  There is no amount of money that is commercially reasonable. Once an ISP is allowed to charge providers for faster or better access to users, ISPs will become the arbiters of which businesses succeed on the Internet and which businesses will fail.  The business that will succeed will, of course, be partners of said ISPs and also a few large corporations that can afford the new ISP tolls.  This cannot happen.  This will break the way the Internet has always worked.  It will hinder the free market spirit of the Internet.  It will diminish free speech by hindering one of the most popular methods of communication.

Net neutrality is good for free speech and good for the US economy.  It helps small businesses.  And allowing ISPs to charge content providers for access to Internet consumers is not net neutrality.

There are several things that the FCC can do to make sure the Internet stays open.  First and foremost, the FCC must classify broadband Internet as a telecommunications service (i.e. a common carrier).  This will give the FCC power to regulate ISPs like other common carriers.  Secondly, the FCC should make sure wording includes mobile carriers.  Mobile users are becoming a larger part of total Internet users every day and they may one day become the majority, if they haven't already.  Mobile Internet is part of the Internet.  And, thusly, mobile providers (AT&T, Sprint, T-Mobile, and Verizon included) are also Internet service providers.

Lastly, the most effective way to make sure broadband Internet is keep fast and open is to declare it a public utility, like water and electricity.  This may be difficult to square with mobile providers, but it can certainly be done with cable and DSL ISPs.  Just like electricity and water, broadband Internet is expensive to rollout to neighborhoods and costly to maintain.  Most areas only have one or two providers because of this natural monopoly.  Fast Internet service has become a necessity in this day and age, becoming one of the most efficient ways to do business, communicate with friends, and disseminate information.

Let's keep the Internet free and open for everyone.  Thank you for listening.

Sincerely,
Omari Christian

Friday, May 9, 2014

Tech Support: Step-By-Step Procedure for Removing Viruses and Other Malware

So, you've got a virus. Or maybe you don't. All you know is you have a Windows PC and the damned thing isn't working like it should. I’ve had this issue countless times, and approximately twice I ended up having actual malware. There’s a few ways to proceed in a situation like this.

You can call tech support for help, but that costs money and takes time. Or you can download and install Trend Micro HijackThis, run it, and post the log on some computer tech support forum. That seems to work for a lot of people, but requires finding the right forum and waiting for kind computer geeks to solve your problem.

Finally, you can try to get rid of the malware yourself using a good malware removal tool. In this article, I’ll describe this option, taking you through my usual process of diagnosing and removing a computer malware from a computer.

"Malware" is any malicious software, including worms, trojans, and viruses. "Anti-virus" is used to refer to any software that can get rid of malware. For a complete dictionary of malware types, check out Viruses, Spyware, and Malware: What's the Difference?

Diagnosing How Bad It Is
First, turn on the computer. If the computer won't turn on--like nothing shows up on the monitor--then you're kinda hosed. It probably wasn't malware. Malware that can damage hardware or wreck your firmware (i.e. your motherboard BIOS) is rare. It's likely that your hardware is broken for other reasons, like being super old or you dropping your computer down a flight of stairs. Seek help elsewhere.

If the computer turns on, but won’t boot into Windows. Several things could be wrong. It could be your hardware, or perhaps malware just corrupted your Windows installation. In any case, that’s more advanced than what we’ll try to solve here.

If the computer turns on and boots into Windows, but you have weird problems there--popups asking you to buy something to “fix” your computer, your browser has been changed to a spammy homepage, your anti-virus won’t run, etc.--then you might have typical malware.

Removing the Malware Manually
When you start up your computer, even before you open your web browser, do you get popups telling you how to fix your computer or speed up your computer? Did you install the program that’s telling you this?

If you have some unknown program that’s telling you that it can fix your computer, you have malware. This is the most common malware I’ve seen in the last few years. One example I recently uninstalled from a friend’s computer was a piece of malware called PC Fix Speed.

image from malwaretips.com

PC Fix Speed and similar fake anti-virus programs want you to be concerned about fake issues with your computer so that you’ll pay them for the full version of their software or for customer support. This type of malware is called "scareware" because the intent is to scare you into paying for their services. God knows what happens after you pay them (is the “fix” a program that uninstalls their own malware?).

Getting the Fix
Download Malwarebytes Anti-Malware. This program’s preventive anti-virus capabilities may be mediocre, but its ability to find and remove existing malware is unparalleled. In the several years I've been using it, I've found no better product for removing existing malware on a computer.

If downloading the software using Internet Explorer doesn’t work, try a different browser. Once I had to deal with malware that had hijacked Internet Explorer, making visiting any website impossible. Firefox and Chrome are a bit more secure.

If you still can’t download Malwarebytes Anti-Malware for whatever reason, you’re going to have to use a different computer to download it. Transfer the installation file (usually named mbam-setup.exe or something similar) to a USB flash drive or external hard drive and then use that to transfer it to your malware-ridden computer.

Closing the Gates
Once I no longer need the Internet, I unplug network cables from my computer and turn off the computer’s wifi, if it has wifi, in case the malicious program is spyware. Spyware can upload files from your computer so that hackers can obtain your passwords, credit card numbers and other personal data that can be used for nefarious deeds. Turning off your Internet prevents this.

Installation
Install Malwarebytes Anti-Malware. If you’re really unlucky, the malware has hijacked your registry and taken control of executable files. Installation will not work, and instead a different program will open, or the installation file won’t open at all. This happened to me once. I couldn’t open any programs at all except for the malware.

Getting Control of Your Executable Files (If Needed)
1. Boot into Safe Mode (Directions) and log in to your account or the Administrator account.
2. Click the Start button and type regedit in the Search box.
3. Right-click Regedit.exe in the returned list and click Run as administrator.
4. If a popup asks you if you want the program to be able to make changes to your computer, click Yes.
5. Browse to the following registry key: HKEY_CLASSES_ROOT\.exe
6. With .exe selected, right-click (Default) and click Modify…
7. Change the Value data to exefile.
modify registry.jpg


8. Browse to and then click on the following registry key: HKEY_CLASSES_ROOT\exefile
9. With exefile selected, right-click (Default) and click Modify…
10. Change the Value data to "%1" %*
(That's quotation marks, percent sign, one, quotation marks, space, percent, asterisk.)
11. Browse to and then click on the following registry key: KEY_CLASSES_ROOT\exefile\shell\open
12. With open selected, right-click (Default) and click Modify…
13. Change the Value data to "%1" %*
14. Close the Registry Editor and restart your PC.

Directions edited from Microsoft Support.

Running Anti-Malware
Malwarebytes Anti-Malware
Once Malwarebytes Anti-Malware is installed, run it by clicking Scan Now in the bottom-right corner. Scanning takes a while. Select and remove all malware that Anti-Malware finds.

Install an Anti-Virus
Now your malware is (hopefully) gone. If you don't have a anti-virus program already, install one to prevent this from occurring again. I use Microsoft Security Essentials, but I've used Avast and AVG in the past. If you don't mind paying some money, McAfee and Norton make anti-virus programs that are probably better than the free options I listed.

Good luck!

Saturday, April 26, 2014

Policy: How the FCC is Killing the Internet and How to Stop It

For a better understanding of net neutrality see my article Net Neutrality 102.

On April 23, the Federal Communications Commission (FCC) said "that it would propose new rules that allow companies like Disney, Google or Netflix to pay Internet service providers (ISPs) like Comcast and Verizon for special, faster lanes to send video and other content to their customers." (NY Times)  Many people, including myself, see this as blatant violation of net neutrality principles. For example, former FCC commissioner Michael Copps, Sen. Al Franken, and Sen. Cory Booker are concerned about the proposed rules.  Senator Bernie Sanders said, "Under this terribly misguided proposal, the Internet as we have come to know it would cease to exist and the average American would be the big loser."

FCC Chairman Tom Wheeler denies that the proposed rules would gut net neutrality.  But as Sean Hollister of The Verge says:
The problem, which Wheeler's statement doesn't refute, is that the FCC intends to say that it's okay to discriminate against traffic if content providers don't pay the ISPs a "commercially reasonable" fee. While the FCC chairman says that "behavior that harms consumers or competition will not be permitted," any fee might risk harming both, even if it's tiny.
How did we get here?  Why is the FCC proposing new rules anyway?  Let's take a step back to explain the situation.

A Quick History of Net Neutrality in the US
Since the beginning of the Internet, ISPs have been net neutral.  Data was data, ISPs didn't discriminate against different types of traffic, and any small personal website could be accessed as quickly as a large corporation's website.  However, more recently ISPs have occasionally blocked or throttled different websites and types of traffic for various reasons.  The recent deal between Netflix and Comcast, was definitely a money grab by Comcast that was arguably non-net neutral.

The FCC, which regulates things like TV and radio, occasionally punished ISPs that violated net neutrality by fining them.  However, in January a federal appeals court stated that the FCC technically didn't have the authority to regulate ISPs. The FCC classifies broadband Internet as an "information service," which limits the FCC's power to regulate broadband ISPs.  Instead of reclassifying broadband Internet as a "telecommunications service"--which would give the FCC more power to regulate them--or appealing the court's decision, Wheeler has decided to rewrite its net neutrality rules.

It now appears that the new rules are not net neutral, throwing the Internet into jeopardy for all Americans.

What Could Happen
I, along with others, have always suspected that if and when ISPs try to kill net neutrality, it would be by charging users for access or for speedy access to certain types of Internet content.  For example, Time Warner Cable might charge Internet subscribers an extra $3 per month to access netflix.com.

That isn't how it's happening.  Instead, ISPs are charging the websites for access to subscribers.  First, AT&T announces "Sponsored Data", which sounds nice but is just a way to block certain websites once a subscriber's data cap is reached.  As Nilay Patel from The Verge says, "that's not fair competition, that's just pay to play."  Next, Comcast throttles Netflix data unless Netflix agrees to pay up.  Now, the FCC announces new rules that imply that this is just the beginning.

In the future, ISPs may charge all websites a fee for access to their users, possibly in speed tiers.  Websites and services that don't pay up will be unreachable by users.  This may sound like perfectly fine, net-neutral behavior to some, like FCC Chair Wheeler, because users cannot pay a fee to get access to these websites.  However, the result is the same: small websites for small businesses, small organizations and non-wealthy individuals will have less of a chance of being seen than large websites run by corporations that can afford these new Internet tolls.

(This is also a reason to oppose the Comcast-Time Warner Cable merger.  While the two companies do not compete for ISP customers because they are in different markets, they will compete for content provider customers.  For example, if Netflix opted not to pay Comcast a fee for normal streaming speeds, they'd only be throttled for half of the country's broadband users.  After the merger they'd be throttled for the whole country.  Comcast will control most American's access to fast Internet and can charge content providers whatever they want.)

Why Is the FCC Trying to Kill Net Neutrality?
Why is this happening?  It's hard to say for sure.  Obama promised net neutrality laws.  Then he strangely nominated Tom Wheeler for FCC Chair.  Wheeler, who lobbied to deregulate the cable industry in the 80's, and lobbied for the wireless industry later.  He supported the T-Mobile and AT&T merger that thankfully failed.  But, when appointed, Wheeler said, "we're pro-open networks," giving hope to net neutrality supporters everywhere.  It appears that hope was misplaced.

Obama's nomination of Wheeler doesn't make much sense, but Wheeler's support for deregulation and supporting huge, oligopolistic corporations--especially the wireless telecoms--is well-established.  Wheeler is just supporting his friends and ex-coworkers.

What Can We Do?
We can still save the Internet.  Contact your representative, your senators, and FCC Chair Tom Wheeler (tom.wheeler@fcc.gov) and tell them to oppose any new FCC rules that aren't net neutral.  Tell them to get the FCC to reclassify broadband providers as "telecommunications services" instead of "information services" so that the FCC can continue using net neutral rules already in place.  If enough people make their voice heard, we can keep the Internet free and open.

Saturday, March 22, 2014

Code: Why Google Guice is Evil

Google Guice (pronounced like "juice") is a dependency injection framework for Java. We use it quite often at work. It’s awful. I think the evils of Guice can best be explained by relating a true story as a second person narrative:
You have a huge project in a huge codebase and your team is using Guice. You notice an instance variable of some type (say, InitHelper) in some class, and the variable is injected into the constructor. You want to find out where that variable is initialized so you can tell what it's settings are. Sure, you can port the whole project into your IDE, compile it and run it with some fake arguments for the backends (or take an hour to initialize the real backends), pause the debugger and see what what the settings were for the that variable, but that would take a long time and you feel that you shouldn't have to run and debug code just to see what it is doing.

Luckily, you work for a top tier company with a searchable code repository. You look for all instances where the class constructor is called.

Nowhere. The constructor is called nowhere (except the tests) because your team is using Guice. No matter. You just need to look in the module. Your team had the foresight to place the module that binds the class in the same directory as the class. "The day is mine!" you exclaim, perhaps a bit prematurely.

Huh. It's not bound in the module. That's unfortunate, but just one of the side effects of being Guicy: sometimes your teammates inject variables from different directories. No biggie, just see who's using your module. Surely there's an @Inject or @Provides that will initialize your variable in another module that includes your module.

Everyone. There must be 100 classes using that module. Is that Fortran code from 1970s that uses the module? Okay, you're going about this the wrong way. You know where the main class of your binary is. Look at the main class and work from there. Is the variable bound there?

No, it's not. Because that would be easy. If it were easy, everyone would do it. And we can't have that. Instead, the main class must include a module that binds the variable you're looking for. Just look through the module(s) included in the main class. How many could there possibly be?

All of them. All of the modules ever written since time immemorial are included in the getModules() method of the main class. Well, adversity builds character. What a great person you're going to be by the end of this! Time to get super serious and just search the repository for "bind(InitHelper.class)".

You found it! The class is bound in a module... in some other team's code you've never seen before. And as far as you can tell, your team doesn't use it. You could grab your best hunting dog and follow the trail to see if something that uses this module is used in your code, or you save a lot of time not following what is probably a dead end. You go with the latter.

Look, you're smarter than this. And can do multiline searches in your repository. If it's not bound manually with a call to bind() then it's provided with a provider, right? Just search for "@Provides" followed by an optional "@Singleton" followed by a newline followed by a variable amount of space followed by the variable's class name. By the power of regex!

Quite a few modules. Let's narrow that search to your team's directory.

Nothing. That's fine, you didn't really want to find it anyway.


If that story didn’t do anything for you, let me explain it in more explicit terms. Guice is awful, not just because it turns your improper type casts from compile-time errors to runtime errors, but also because it can turn your code into spaghetti code, where searching for a simple injected variable becomes a quest to find a single needle in a silver haystack.

Disclaimer: Obviously, a code framework cannot be “evil.” And Guice may not always be necessarily bad. When used properly, I'm sure you can achieve Guice’s benefits while still writing maintainable, easy-to-read code. I’m just not quite sure what those benefits are. I've certainly never seen any code where it made things easier to read or to write.

An example of the futility of Guice can be found on the Motivation page of the Guice User’s Guide [link updated; the new Motivation page doesn't have a comments section].  The Motivation page describes a situation where Guice is used to make swapping test components and real components easier in a billing service.  That page ends with the following example of Guice injection:

Injector injector = Guice.createInjector(new BillingModule());
BillingService billingService = injector.getInstance(BillingService.class);

In the comments section, one person asks a good question:
As a cynic of [dependency] injection frameworks, can someone tell me what the advantage is of

Injector injector = Guice.createInjector(new BillingModule());
BillingService billingService = injector.getInstance(BillingService.class);

over

BillingService billingService = new BillingService(new BillingModule());

please?
He (or she) has a point. And in my opinion, his question is never answered satisfactorily. Why depend on annotations and reflection (the GOTOs of the Java language) when you can use the perfectly legitimate new operator? In his example, BillingService gets all its instance variables from the module argument. The constructor for BillingService would look something like this:

public RealBillingService(BillingModuleInterface realOrFakeBillingModule) {
    this.processor = realOrFakeBillingModule.getProcessor();
    this.transactionLog = realOrFakeBillingModule.getTransactionLog();
}

In this case, BillingModuleInterface is an interface for a module that may contain real or mocked-out billing components. It’s easy to read because it’s using good old fashioned typical Java. It’s easy to write because any missed dependencies or casting errors will be found at compile time. It’s easy to add a new module because you can simply create a new class that extends BillingModuleInterface.

In conclusion, the Guice framework is a poor choice for most, if not all, projects. It’s especially bad for large projects in large codebases where finding an injected dependency may be a Sisyphean ordeal.